Search PubMed⌕ Search

Biomedical subjects

D R Forsdyke

Publications and source records attributed to D R Forsdyke.

At least 19 recordsLinked to original sources

Low-complexity segments in Plasmodium falciparum proteins are primarily nucleic acid level adaptations.

Protein segments that contain few of the possible 20 amino acids, sometimes in tandem repeat arrays, are referred to as containing "simple" or "low-complexity" sequence. Many Plasmodium falciparum proteins are longer than their homologs in other species by virtue of their content of such low-complexity segments that have no known function; these are interspersed among segments of higher complexity to which function can often be ascribed. If there is low complexity at the protein level, there is likely to be low complexity at the corresponding nucleic acid level (departure from equifrequency of the four bases). Thus, low complexity may have been selected primarily at the nucleic acid level and low complexity at the protein level may be secondary. In this case, the amino acid composition of low-complexity segments should be more reflective than that of high complexity segments on forces operating at the nucleic acid level, which include GC-pressure and AG-pressure. Consistent with this, for amino acid determining first and second codon positions, open reading frames containing low-complexity segments show increased contributions to downward GC-pressure (revealed as decreased percentage of G+C) and to upward AG-pressure (revealed as increased percentage A+G). When not countermanded by high contributions to AG-pressure, low-complexity segments can contribute to base order-dependent fold potential; in this respect, they resemble introns. Thus, in P. falciparum, low-complexity segments appear as adaptations primarily serving nucleic acid level functions.

Amino Acids↗

Symmetry observations in long nucleotide sequences: a commentary on the Discovery Note of Qi and Cuticchia.

The relative quantities of bases in DNA were determined chemically many years before sequencing technologies permitted direct counting of bases. Apparently unaware of the rich literature on the topic, bioinformaticists are today rediscovering the 'wheels' of Chargaff, Wyatt and other biochemists. It follows from Chargaff's second parity rule (%A = %T, %G = %C for single stranded DNA) that the symmetries observed for the two pairs of complementary mononucleotide bases, should also apply to the eight pairs of complementary dinucleotide bases, the thirty-two pairs of complementary trinucleotide bases, etc. This was made explicit by Prabhu in 1993 in a study of complete genomes and long genome segments from a wide range of taxa, and was rediscovered by Qi and Cuticchia in 2001 in a study of complete genomes. It follows from Chargaff's GC-rule (%GC tends to be uniform and species specific) that, within a species, oligonucleotides of the same GC% will be at approximately equal quantities in single stranded DNA. Thus, for example, while quantities of CAT and ATG (reverse complements) will be closely correlated because of both of the above Chargaff rules, CAT and GTA (forward complements) will show some correlation only because of the latter rule. The need for complete genomic sequences in bioinformatic analyses may have been somewhat overplayed.

Base Sequence↗

Adaptive value of polymorphism in intracellular self/not-self discrimination?

A microbial pathogen species can adapt to its host species to the extent that members of the host species are uniform. Loss of this uniformity would make it difficult for a pathogen species to transfer, from one member of the host species to another, what it had "learned" through selection of its members with advantageous mutations. The existence of major histocompatibility complex (MHC) polymorphism indicates that non-uniformity within a species is an effective host defence strategy. By virtue of this molecular discontinuity among its members the host species can "present a moving target" to the pathogen. Many proteins other than MHC proteins show polymorphism - a phenomenon which has suggested that mutations in regions of protein molecules which do not affect overt function are neutral. However, in the context of the author's differential aggregation theory of intracellular self/not-self discrimination as previously applied to the problem of the antigenicity of cancer cells, such polymorphism should serve for the recruitment of subsets of self-antigens into the antigenic repertoire of an infected cell. These would act as "intracellular antibodies" by virtue of their weak, but specific, aggregation with pathogen proteins. Peptides from the self-antigens, as well as (or instead of) those from the antigens of the pathogen, would then serve as targets for attack by cytotoxic T cells. Thus, polymorphism of intracellular proteins should be of adaptive value, serving to amplify and individualize the immune response to intracellular pathogens.

Animals↗

Introns resolve the conflict between base order-dependent stem-loop potential and the encoding of RNA or protein: further evidence from overlapping genes.

Many eukaryotic genes are split into exons and introns, the latter being removed post-transcriptionally so that only exon sequences appear in cytoplasmic RNAs. Since introns appear in both protein-encoding RNAs and non-protein-coding RNAs, they interrupt genetic information per se, not just protein-encoding information. A DNA sequence has the potential to carry more than one type of genetic information, but different types may conflict. Thus, it has been proposed that introns arose because sequences were unable to contain concomitantly complete information for the encoding both of stem-loops and of cytoplasmic products (protein and/or RNA). Stem-loop potential is held to be selectively advantageous since it promotes the recombination-dependent correction of genetic errors. Stem-loop potential, the best local measure of which is base order-dependent stem-loop potential, tends to be less in exons than in introns. This is particularly evident in genes evolving rapidly under positive Darwinian selection, where the protein-encoding function is dominant. Evidence is now presented that the rare regions where genes overlap also impose excessive encoding demands so that the concomitant coding of base order-dependent stem-loop potential is decreased. Our results are consistent with the hypothesis that sequences with high stem-loop potential arose in the early 'RNA world'. Ancestors of modern genes would have entered this world when sequences (exons) encoding cytoplasmic products, were interspersed with sequences (introns) encoding selectively advantageous stem-loops. Purine-loading pressure would also have favoured intron formation.

Animals↗

Double-stranded RNA as a not-self alarm signal: to evade, most viruses purine-load their RNAs, but some (HTLV-1, Epstein-Barr) pyrimidine-load.

For double-stranded RNA (dsRNA) to signal the presence of foreign (non-self) nucleic acid, self-RNA-self-RNA interactions should be minimized. Indeed, self-RNAs appear to have been fine-tuned over evolutionary time by the introduction of purines in clusters in the loop regions of stem-loop structures. This adaptation should militate against the "kissing" interactions which initiate formation of dsRNA. Our analyses of virus base compositions suggest that, to avoid triggering the host cell's dsRNA surveillance mechanism, most viruses purine-load their RNAs to resemble host RNAs ("stealth" strategy). However, some GC-rich latent viruses (HTLV-1, EBV) pyrimidine-load their RNAs. It is suggested that when virus production begins, these RNAs suddenly increase in concentration and impair host mRNA function by virtue of an excess of complementary "kissing" interactions ("surprise" strategy). Remarkably, the only mRNA expressed in the most fundamental form of EBV latency (the "EBNA-1 program") is purine-loaded. This apparent stealth strategy is reinforced by a simple sequence repeat which prefers purine-rich codons. During latent infection the EBNA-1 protein may evade recognition by cytotoxic T-cells, not by virtue of containing a simple sequence amino acid repeat as has been proposed, but by virtue of the encoding mRNA being purine-loaded to prevent interactions with host RNAs of either genic or non-genic origin.

Animals↗

Chargaff's legacy.

Of Chargaff's four rules on DNA base composition, only his first parity rule was incorporated into mainstream biology as the DNA double helix. Now, the cluster rule, the second parity rule, and the GC rule, reveal the multiple levels of information in our genomes and potential conflicts between them. In these terms we can understand how double-stranded RNA became an intracellular alarm signal, how potentially recombining nucleic acids can distinguish between 'self' and 'not-self' so leading to the origin of species, how isochores evolved to facilitate gene duplication, and how unlikely it is that any mutation can ever remain truly neutral.

Animals↗

Haldane's rule: hybrid sterility affects the heterogametic sex first because sexual differentiation is on the path to species differentiation.

Prevention of recombination is needed to preserve both phenotypic differentiation between species and sexual phenotypic differentiation within species. For species differentiation (speciation), isolating barriers preventing recombination may be pre-zygotic (gamete transfer barriers), or post-zygotic (either a developmental barrier resulting in hybrid inviability, or a chromosomal-pairing barrier resulting in hybrid sterility). The sterility barrier is usually the first to appear and, although often initially only manifest in the heterogametic sex (Haldane's rule), is finally manifest in both sexes. For sexual differentiation, the first and only barrier is chromosomal-pairing, and always applies to the heterogametic sex. For regions of sex chromosomes affecting sexual differentiation there must be something analogous to the process generating the hybrid sterility seen when allied species cross. Explanations for Haldane's rule have generally assumed that the chromosomal-pairing barrier initiating evolutionary divergence into species is due to incompatibilities between gene products ("genic), or sets of gene products ("polygenic), rather than between chromosomes per se ("chromosomal"). However, if chromosomal incompatibilities promoting incipient sexual differentiation could also contribute to the process of incipient speciation, then a step towards speciation would have been taken in the heterogametic sex. Thus, incipient speciation, manifest as hybrid sterility when "varieties" are crossed, would appear at the earliest stage in the heterogametic sex, even in genera with homomorphic sex chromosomes (Haldane's rule for hybrid sterility). In contrast, it has been proposed that Haldane's rule for hybrid inviability needs differences in dosage compensation, so could not apply to genera with homomorphic sex chromosomes.

Animals↗

Crossover hot-spot instigator (Chi) sequences in Escherichia coli occupy distinct recombination/transcription islands.

Crossover hot-spot instigator (Chi) sequences (5'-GCTGGTGG-3') are orientation-dependent, strand-specific sequences implicated in RecA-mediated DNA recombination. In Escherichia coli and Haemophilus influenzae Chi and Chi-like sequences preferentially locate to approx. 1kb recombination 'islands' in the mRNA-synonymous strands of open reading frames (ORFs). Since mRNA-synonymous strands follow Szybalski's transcription direction rule in being G-rich, and the average ORF is about 1kb, then, on this basis alone, Chi sequences are seen to reside in 1kb G-rich 'islands'. However, RecA preferentially binds GT-rich sequences, suggesting that genomic context might potentiate Chi action. Consistent with this, we report for E. coli that 1kb sequence windows with Chi near their centres are a distinct subset of total 1kb windows, the mRNA-synonymous strands being preferentially enriched in both G and T. Chi function might be particularly important for bacteria that survive high temperature and radiation. These often exist in habitats where recombination with E. coli DNA would be unlikely, so canonical Chi sequences might not confer a selective disadvantage in this respect. In general, Chi sequences are not more frequent in thermophilic bacteria and Deinococcus radiodurans, than in E. coli and other mesophilic bacteria. Only two of five thermophilic bacteria examined showed preferential location of Chi sequences to mRNA-synonymous strands. In the thermophile Methanococcus jannaschii, windows containing the canonical Chi sequence do not form a distinct subset. We suggest that in thermophilic bacteria and D. radiodurans the Chi function may be achieved by sequences that differ from the canonical Chi sequence, or that the number of these sequences is sufficient, or that the Chi function is unnecessary.

Base Composition↗

Thermophilic bacteria strictly obey Szybalski's transcription direction rule and politely purine-load RNAs with both adenine and guanine.

When transcription is to the right of the promoter, the "top," mRNA-synonymous strand of DNA tends to be purine-rich. When transcription is to the left of the promoter, the top, mRNA-template strand tends to be pyrimidine-rich. This transcription-direction rule suggests that there has been an evolutionary selection pressure for the purine-loading of RNAs. The politeness hypothesis states that purine-loading prevents distracting RNA-RNA interactions and excessive formation of double-stranded RNA, which might trigger various intracellular alarms. Because RNA-RNA interactions have a distinct entropy-driven component, the pressure for the evolution of purine-loading might be greater in organisms living at high temperatures. In support of this, we find that Chargaff differences (a measure of purine-loading) are greater in thermophiles than in nonthermophiles and extend to both purine bases. In thermophiles the pressure to purine-load affects codon choice, indicating that some features of their amino acid composition (e.g., high levels of glutamic acid) might reflect purine-loading pressure (i.e., constraints on mRNA) rather than direct constraints on protein structure and function.

Adenine↗

Two levels of information in DNA: relationship of Romanes' "intrinsic" variability of the reproductive system, and Bateson's "residue" to the species-dependent component of the base composition, (C+G)%.

In 1886 Charles Darwin's research associate George Romanes published a paper entitled "Physiological Selection: An Additional Suggestion on the Origin of Species". This was criticized by his Victorian contemporaries and largely ignored by those who followed. However, the recent recognition of two levels of information in DNA suggests that Romanes had solved the major problems with Darwin's theory. It was apparent from the outset that the form of reproductive isolation likely to apply most generally to initial species divergence (hybrid sterility), would depend on differences, not in "primary" information ("genic"), but in "secondary" information ("chromosomal"). This viewpoint, further elaborated by Bateson & Saunders (1902), White (1978), and King (1993), is criticized by the genic school (Coyne & Orr, 1998) because it requires visible differences between chromosomes, and appears not to explain Haldane's rule. However, chromosomal differentiation with respect to the species-dependent component of base composition [(C+G)%; Forsdyke, 1996] appears to resolve these problems. Because it explained so much, it was easy to believe that the genic viewpoint explained everything. Romanes and Bateson thought otherwise. We are only just beginning to recognize what they were trying to tell us.

Animals↗

Accounting units in DNA.

Chargaff's first parity rule (%A=%T and %G=%C) is explained by the Watson-Crick model for duplex DNA in which complementary base pairs form individual accounting units. Chargaff's second parity rule is that the first rule also applies to single strands of DNA. The limits of accounting units in single strands were examined by moving windows of various sizes along sequences and counting the relative proportions of A and T (the W bases), and of C and G (the S bases). Shuffled sequences account, on average, over shorter regions than the corresponding natural sequence. For an E. coli segment, S base accounting is, on average, contained within a region of 10 kb, whereas W base accounting requires regions in excess of 100 kb. Accounting requires the entire genome (190 kb) in the case of Vaccinia virus, which has an overall "Chargaff difference" of only 0.086% (i.e. only one in 1162 bases does not have a potential pairing partner in the same strand). Among the chromosomes of Saccharomyces cerevisiae, the total Chargaff differences for the W bases and for the S bases are usually correlated. In general, Chargaff differences for a natural sequence and its shuffled counterpart diverge maximally when 1 kb sequence windows are employed. This should be the optimum window size for examining correlations between Chargaff differences and sequence features which have arisen through natural selection. We propose that Chargaff's second parity rule reflects the evolution of genome-wide stem-loop potential as part of short- and long-range accounting processes which work together to sustain the integrity of various levels of information in DNA.

Base Composition↗

Deviations from Chargaff's second parity rule correlate with direction of transcription.

The distribution of deviations from Chargaff's second parity rule was examined for overlapping sequence windows of a length (1 kb) predicted to be suitable for detecting correlations with functional features of DNA. For long genomic segments from E. coli, Saccharomyces cerevisiae, and Vaccinia virus, Chargaff differences for the W bases and/or for the S bases correlate with transcription direction and gene location. For W-rich genomes, the mRNA-synonymous strand contains regions which, if extruded from negatively supercoiled DNA, would fold to generate stem-loop structures with A-rich loops. Similarly, for S-rich genomes the loops would be G-rich. We suggest that the disposition of genes in nucleic acid sequences arises from their having to adapt to a preexisting mosaic of genomic regions, each distinguished by its potential to extrude single-strand loops enriched for a particular base (or two non-Watson-Crick pairing bases). The mosaic would have facilitated the intrastrand and interstrand accounting required for correction of mutations, and would have evolved in the early RNA world before the emergence of protein-encoding capacity. The preexisting mosaic would have determined transcription direction since there is pressure for all mRNAs of a cell to have purine-rich loops, thus decreasing loop-loop interactions which might lead to formation of "self" sense-antisense RNA duplexes.

Animals↗

Heat shock proteins as mediators of aggregation-induced 'danger' signals: implications of the slow evolutionary fine-tuning of sequences for the antigenicity of cancer cells.

Organisms 'tune' to their environment through adaptations which confer a selective advantage. However, in complex systems, a primary change of positive adaptive value might have multiple minor secondary effects, usually of negative adaptive value, which could invoke further counter-adaptations. This 'fine-tuning', a 'debugging', mainly at the intracellular level, would appear an evolutionary burden detracting from the positive nature of the primary change. However, if the primary mutation is in a potential oncogene, secondary, short-term effects may include the recruitment, in an apparently random manner, of unmutated non-oncogene products into the antigenic repertoire of the cancer cell. This 'danger' signal, provided by the co-aggregation of oncogene and non-oncogene products, would be mediated by inducible heat-shock proteins (Hsps), and lead to display of corresponding MHC-peptide complexes. It was argued previously that T cells specific for peptides from most 'self' intracellular antigens are not eliminated during T cell 'education', and so would be available for subsequent immune activation by the corresponding peptides. These considerations might explain why cancer specific antigens have been so elusive, why cancer antigenicity is often individual specific, and why therapeutic approaches involving complexes of peptides with Hsps may be successful.

Antigens, Neoplasm↗

Correlation of chi orientation with transcription indicates a fundamental relationship between recombination and transcription.

Cross-over hot-spot instigator (Chi) sequences (5'-GCTGGTGG-3') are abundant, strand-specific, sequences, which locally increase recombination in Escherichia coli. Located within G-rich 'recombination islands', Chi orientations correlate with the orientations both of DNA replication and of transcription. Consistent with evidence from eukaryotic systems for a fundamental relationship between recombination and transcription, we find for E. coli Chi sequences, and for Haemophilus influenzae Chi-like sequences, that orientations correlate better with transcription than with replication. Complying with Szybalski's transcription direction rule, open reading frames in these prokaryotes have purine-rich mRNA-synonymous DNA strands. Hence, the G-richness of 'recombination islands' may reflect their correspondence with 'transcriptional islands' (genes). Comparison of a natural with the corresponding shuffled sequence, indicates a base order-dependent island unit of approx. 1kb. 1998 Elsevier Science B.V.

Base Composition↗

An alternative way of thinking about stem-loops in DNA. A case study of the human G0S2 gene.

Single strands extruded from duplex DNA have the potential to form stem-loop structures, which may be involved in the homology search preceding recombination. The total stem-loop potential in a sequence window can be analysed in terms of the relative contributions of base composition and base order. There are at least 10 base composition-determined parameters of relevance to the energetics of stem-loop formation. These are the quantities of the four bases themselves, and six derived parameters: ATmin, CGmin, Chargaff differences for the W and S bases, and two base products. The quantities of the least represented base of a Watson-Crick base pair (ATmin, CGmin) might provide an index of the total stem potential of a widow. The degrees to which one base of a Watson-Crick pair exceeds the other (the Chargaff differences for the W bases and for the S bases) might provide an index of the total loop potential of a window. Base products (A x T, C x G) might provide an index both of stem and of loop potentials. Multiple regression analysis of the relationship of the 10 parameters to the energetics of stem-loop formation in the G0S2 gene reveals major roles of S bases, and of base products. While base composition may primarily serve genome or genome sector "strategies", it becomes of local relevance in the case of CpG islands. Base order serves many local "strategies", whose demands may conflict. Base order serves the encoding of protein or of recognition motifs for regulatory factors. On the other hand there appear to be circumstances under which base order synergizes with, or antagonises, base composition in determining total stem-loop potential. Antagonism is evident when the base composition-dependent component of the stem-loop potential of a region is greater than the total stem-loop potential of that region.

DNA↗

Expression and processing of G0/G1 switch gene 24 (G0S24/TIS11/TTP/NUP475) RNA in cultured human blood mononuclear cells.

The human G0/G1 switch (G0S) gene, G0S24, and its rodent immediate-early homolog (TIS11, TTP, NUP475) are part of a mammalian gene family whose members encode CCCH zinc finger domains and domains similar to part of the large subunit of RNA polymerase II and to the Mei2 regulator of G1 arrest in fission yeast. We compared the RNA expression of G0S24 with that of other G0S genes in cultured blood mononuclear cells and examined the levels of various RNA processing intermediates. Freshly isolated cells contained high levels of several G0S RNAs, which declined by 24 h, suggesting transient spontaneous stimulation during cell purification (Heximer et al., 1996). However, in cells preincubated for 24 h, G0S24 RNA levels remained much higher than those of other G0S genes (107+/-42 x 10(6) molecules/microg of RNA); stimulation with lectin (Con-A) further increased G0S24 RNA, much of which remained nuclear. Like those of FOS/G0S7, EGR1/G0S30 and of the gene encoding the regulator of G protein signalling 1 (RGS1), G0S24 RNA levels increased more in response to a protein kinase C activator than to a calcium ionophore, whereas the opposite held for FOSB/G0S3 and RGS2/G0S8. With appropriate PCR primer pairs, we showed a G0S24 RNA processing intermediate, which crossed the exon-1/intron boundary, and nonpolyadenylated nuclear RNA extending into the 3' flank, where there is a second CpG island. The concentration of the latter intermediate (1.2+/-0.2 x 10(6) molecules/microg of RNA), which increased transiently on cell stimulation, did not account for all G0S24 nuclear RNA. The levels of G0S24 RNA and both intermediates were increased by the protein synthesis inhibitor cycloheximide, consistent with regulation by a labile repressor.

Base Sequence↗

The normal copy of the G0S19-3-associated, CpG island-containing, upstream sequence is downstream of G0S19-2/MIP1alpha in association with a TRE17 oncogene.

The G0S19-1/MIP1alpha and G0S19-2/MIP1alpha genes locate to human chromosomes 17q and encode similar copies of the beta-chemokine G0S19/MIP1alpha. The G0S19-3 gene, present in 1 in 4 humans, is a 5' truncated version of G0S19-2; a CpG island-containing upstream sequence (CpG-US), rich in potential transcriptional activation motifs, replaces much of the first intron and the first exon. Sequences hybridizing with the CpG-US sequence, normally exist in all human genomes. Thus, it appears that there has been recombination between a duplicated G0S19 gene and a duplicated CpG-US-like sequence. We have isolated sequences hybridizing with the CpG-US sequence from a human genomic library in bacteriophage lambda. Restriction mapping and sequencing shows a CpG-US-like sequence approximately 8 kb downstream of G0S19-2 (hence, named CpG-DS sequence). The sequence is contiguous with a TRE17 oncogene-associated sequence (GenBank locus HSTRE175). Members of the TRE17 family are known to locate to chromosome 17q (Onno et al., 1993b), and have sequence characteristics suggestive of positive Darwinian selection. Linkage with a TRE17 oncogene may have arisen by recombination and imply no functional relationship. However, it is possible that the CpG-DS may normally regulate TRE17 expression. PCR and sequencing studies indicate the close proximity of other chemokine-related sequences in the 17q11.2 region.

Amino Acid Sequence↗