Search PubMed⌕ Search

Biomedical subjects

Jan Charles Biro

Publications and source records attributed to Jan Charles Biro.

3 recordsLinked to original sources

Indications that "codon boundaries" are physico-chemically defined and that protein-folding information is contained in the redundant exon bases.

BACKGROUND: All the information necessary for protein folding is supposed to be present in the amino acid sequence. It is still not possible to provide specific ab initio structure predictions by bioinformatical methods. It is suspected that additional folding information is present in protein coding nucleic acid sequences, but this is not represented by the known genetic code. RESULTS: Nucleic acid subsequences comprising the 1st and/or 3rd codon residues in mRNAs express significantly higher free folding energy (FFE) than the subsequence containing only the 2nd residues (p < 0.0001, n = 81). This periodic FFE difference is not present in introns. It is therefore a specific physico-chemical characteristic of coding sequences and might contribute to unambiguous definition of codon boundaries during translation. The FFEs of the 1st and 3rd residues are additive, which suggests that these residues contain a significant number of complementary bases and that may contribute to selection for local RNA secondary structures in coding regions. This periodic, codon-related structure-formation of mRNAs indicates a connection between the structures of exons and the corresponding (translated) proteins. The folding energy dot plots of RNAs and the residue contact maps of the coded proteins are indeed similar. Residue contact statistics using 81 different protein structures confirmed that amino acids that are coded by partially reverse and complementary codons (Watson-Crick (WC) base pairs at the 1st and 3rd codon positions and translated in reverse orientation) are preferentially co-located in protein structures. CONCLUSION: Exons are distinguished from introns, and codon boundaries are physico-chemically defined, by periodically distributed FFE differences between codon positions. There is a selection for local RNA secondary structures in coding regions and this nucleic acid structure resembles the folding profiles of the coded proteins. The preferentially (specifically) interacting amino acids are coded by partially complementary codons, which strongly supports the connection between mRNA and the corresponding protein structures and indicates that there is protein folding information in nucleic acids that is not present in the genetic code. This might suggest an additional explanation of codon redundancy.

Amino Acid Sequence↗

Overlapping translation of nucleic acid sequences for bioinformatics applications.

SUMMARY: An alternative method to TblastX has been developed. Nucleic acids in database and query sequences were translated into overlapping protein-like sequences (overlappingly translated sequences or OTSs) before searching with BlastP. Thus, each nucleic acid sequences is represented by a single 'protein like' sequence instead of three 'proteins' in different reading frames. The 3x3 comparison of TblastX is represented by a single comparison, giving faster results. Additional advantages are: (1) it can be more sensitive to detect weak sequence similarities than either blastN or TblastX; (2) codon redundancy is eliminated; (3) the sensitivity to single nucleotide polymorphism, mutation and sequencing errors is reduced; (4) it is insensitive to frame shifts. RESULTS: BlastP using OTS detected about two thirds of blastN and TblastX matches but discovered additional similarities. When blastN and TblastX against nucleic acids were compared to blastP against OTS, identical matches discovered by blastP were generally longer (602, respectively. 213 letters, p<0.01), had higher scores (748 respectively 460 bits, p<0.05) and lower E values (3.16E-20 vs. 1.17E+03, p<0.01) but the percentage identity was lower (25% respectively 61%, p<0.001). A qualitative evaluation with LALIGN showed an improvement of the visualization when OTS-s were used instead of nucleic acids. Many extensive sequence similarities became better visible, for example the repeating similarity between prion protein and human insulin gene micro-satellite, and the surprising similarity between the first part of prion protein coding region and the human pro-insulin (34.4% identity and additional 17.2% similarity through 238 residues, score >295 which is expected 4.6e-18 times by chance).

Amino Acid Sequence↗

Speculations about alternative DNA structures.

An alternative model to the Watson & Crick (W&C) double DNA-spiral and the Pauling & Corey (P&C) triple spiral is presented. In this model: (1). the rotation axis of the polynucleotide chain is in the ribose ring; (2). there is a -H- bond or direct covalent bond between the O2 (PO(4) and C2(') (in ribose) which makes the nucleic acid strands 'stiff'; (3). when there is a covalent bond between O2 and C2('), the unit of the DNA is the ribonucleoside 2('), 3(')-cyclic monophosphate, an intermediate form between DNA and RNA; (4). the bases point outwards from the rotation axis and may interact with each other to connect 2-4 strands together through complementary base pairs; (5). two strands may, but do not necessarily, form a helical structure and if they do, the interacting strands do not turn around each other. The architecture of this model, termed the Homulus DNA model is open (in contrast to the inverted W&C model) and using it might help us to understand the nature of some specific DNA-protein interactions, ordered chromatin formation (coiling and de-coiling), specific gene-to-gene interaction (gene targeting). It is possible that a small portion of the total DNA, the transcribed, 'working DNA', might be built by this way.

Animals↗