Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “mRNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Natural selection and algorithmic design of mRNA.

Messenger RNA (mRNA) sequences serve as templates for proteins according to the triplet code, in which each of the 4(3) = 64 different codons (sequences of three consecutive nucleotide bases) in RNA either terminate transcription or map to one of the 20 different amino acids (or residues) which build up proteins. Because there are more codons than residues, there is inherent redundancy in the coding. Certain residues (e.g., tryptophan) have only a single corresponding codon, while other residues (e.g., arginine) have as many as six corresponding codons. This freedom implies that the number of possible RNA sequences coding for a given protein grows exponentially in the length of the protein. Thus nature has wide latitude to select among mRNA sequences which are informationally equivalent, but structurally and energetically divergent. In this paper, we explore how nature takes advantage of this freedom and how to algorithmically design structures more energetically favorable than have been built through natural selection. In particular: (1) Natural Selection--we perform the first large-scale computational experiment comparing the stability of mRNA sequences from a variety of organisms to random synonymous sequences which respect the codon preferences of the organism. This experiment was conducted on over 27,000 sequences from 34 microbial species with 36 genomic structures. We provide evidence that in all genomic structures highly stable sequences are disproportionately abundant, and in 19 of 36 cases highly unstable sequences are disproportionately abundant. This suggests that the stability of mRNA sequences is subject to natural selection. (2) Artificial Selection--motivated by these biological results, we examine the algorithmic problem of designing the most stable and unstable mRNA sequences which code for a target protein. We give a polynomial-time dynamic programming solution to the most stable sequence problem (MSSP), which is asymptotically no more complex than secondary structure prediction. We show that the corresponding least stable sequence problem (LSSP) is NP-complete, and develop two heuristics for the construction of such sequences. We have implemented these algorithms, and present experimental results placing the high/low stability sequences in context with both wildtype and random encodings. Our implementation has already been applied to the design of RNA "code-words" creating little or no secondary structure in RNA computing (Brenneman and Condon, 2001; Marathe et al., 2001), and we anticipate a variety of other applications of this work to sequence design problems (Skiena, 2001).

Algorithms↗

Determination of the capped site sequence of mRNA based on the detection of cap-dependent nucleotide addition using an anchor ligation method.

The sequence analysis of the 5' ends of cDNAs prepared using the anchor ligation method has revealed that most of the full-length cDNAs have an additional dGMP at their 5' end that is absent in the corresponding genome sequence. Using model RNA transcripts with cap analogues possessing 7-methylguanosine and adenosine, the base of the added nucleotide has been shown to be complementary to the base of the cap analogue, suggesting that the cDNAs possessing an additional dGMP are derived from intact mRNAs with the cap structure. On the other hand, cap-free RNA did not produce cDNA with an extra dGMP. These findings suggest that we can determine whether or not the cDNA starts from the capped site sequence of mRNA based on the presence or absence of an additional dGMP at the 5' end of the cDNA synthesized using the anchor ligation method. This approach will be useful to determine the capped site sequence of mRNA, thus, to identify transcription start sites.

DNA, Complementary↗

Statistical evidence for conserved, local secondary structure in the coding regions of eukaryotic mRNAs and pre-mRNAs.

Owing to the degeneracy of the genetic code, protein-coding regions of mRNA sequences can harbour more than only amino acid information. We search the mRNA sequences of 11 human protein-coding genes for evolutionarily conserved secondary structure elements using RNA-Decoder, a comparative secondary structure prediction program that is capable of explicitly taking the known protein-coding context of the mRNA sequences into account. We detect well-defined, conserved RNA secondary structure elements in the coding regions of the mRNA sequences and show that base-paired codons strongly correlate with sparse codons. We also investigate the role of repetitive elements in the formation of secondary structure and explain the use of alternate start codons in the caveolin-1 gene by a conserved secondary structure element overlapping the nominal start codon. We discuss the functional roles of our novel findings in regulating the gene expression on mRNA level. We also investigate the role of secondary structure on the correct splicing of the human CFTR gene. We study the wild-type version of the pre-mRNA as well as 29 variants with synonymous mutations in exon 12. By comparing our predicted secondary structures to the experimentally determined splicing efficiencies, we find with weak statistical significance that pre-mRNAs with high-splicing efficiencies have different predicted secondary structures than pre-mRNAs with low-splicing efficiencies.

Animals↗

Fusion protein of the paramyxovirus simian virus 5: nucleotide sequence of mRNA predicts a highly hydrophobic glycoprotein.

The nucleotide sequence of the mRNA coding for the fusion glycoprotein (F) of the paramyxovirus, simian virus 5, has been obtained. There is a single large open reading frame on the mRNA that encodes a protein of 529 amino acids with a molecular weight of 56,531. The proteolytic cleavage/activation site of F, to yield F2 and F1, contains five arginine residues. Six potential glycosylation sites were identified in the protein, two on F2 and four on F1. The deduced amino acid sequence indicates that F is extensively hydrophobic over the length of the polypeptide chain. Three regions are very hydrophobic and could interact directly with membranes: these are the NH2-terminal putative signal peptide, the COOH-terminal putative membrane anchorage domain, and the NH2-terminal region of F1.

Amino Acid Sequence↗

Characterization of a sheep elastin cDNA clone containing translated sequences.

mRNA, isolated from the ligamentum nuchae of fetal sheep by guanidine HCl extraction and oligo(dT) cellulose chromatography, was used to synthesize blunt-ended cDNA molecules by the successive application of AMV reverse transcriptase, DNA polymerase and S1 nuclease. The cDNA was centrifuged on a 15-30% sucrose gradient and molecules greater than 700 bp were tailed with dCTP and cloned into the PstI site of pBR322 which had been tailed with dGTP. Ampicillin-sensitive and tetracycline-resistant colonies were screened by in situ hybridization with elastin-enriched mRNA that had been terminally labeled with 32p. Recombinant plasmids prepared from strongly hybridizing colonies were characterized by restriction mapping and the plasmid with the largest insert (1300 bp) thought to contain elastin sequences was characterized in more detail. The nick-translated cDNA hybridized to a single 3.5 kb mRNA species upon blot hybridization, a size identical to that previously identified for chick elastin mRNA (Burnett et al. (1982) J. Biol. Chem. 259, 1569-1572). Nucleotide sequencing of the 5' end of the cDNA demonstrated a sequence which was extremely GC rich and which corresponded to an amino acid sequence partially homologous to that previously identified in porcine tropoelastin (Foster et al. (1973) J. Biol. Chem. 248, 2876-2879). This is the first report of the identification of a plasmid containing sequences complementary to a translated region of elastin mRNA.

Amino Acid Sequence↗

Natural and synthetic heat shock protein gene promoters assayed in Drosophila cells.

Hybrid genes containing mRNA encoding sequences for herpes virus thymidine kinase (tk), chloramphenicol acetyltransferase (CAT), or Drosophila alcohol dehydrogenase (Adh), ligated to truncated Drosophila melanogaster heat-shock protein 70 (hsp 70) gene promoters or to synthetic sequences containing one or several copies of a previously defined heat-shock consensus sequence, were transfected into cultured Drosophila line S3 cells. Each construction was then assayed for gene expression at 25 degrees C and 37 degrees C, using a CAT enzyme assay, slot blot hybridization, or S1 nuclease protection analysis. In the Drosophila cell transient expression assay system, we found that deletions extending beyond position -97, or synthetic constructions containing a single heat shock consensus sequence, were not induced by high-temperature shock. In constructions containing deletions extending to position -186, -130, or -97, in the hsp 70 promoter, and in synthetic constructions containing tandemly spaced heat-shock consensus sequences mRNA transcription was greatly induced by high temperature.

Animals↗

Effect of histones and nonhistone chromosomal proteins on the transcription of histone genes from HeLaS3 cell DNA.

To elucidate the manner in which histones and nonhistone chromosomal proteins interact to render histone genes transcribable in HeLa S3 cells, we have examined transcription of histone mRNA sequences from DNA, as well as from several DNA-chromosomal protein complexes. Histone mRNA sequences were assayed by hybridization to a 3H-labeled single-stranded DNA complementary to histone mRNAs. Our results indicate that DNA is an effective template for transcription of histone mRNA sequences and that histones by themselves inhibit transcription from DNA, including transcription of histone genes, in a dose-dependent, nonspecific manner. When complexed with DNA alone, nonhistone chromosomal proteins do not affect the transcription of histone mRNA sequences. However, when associated with DNA in the presence of histones, nonhistone chromosomal proteins are capable of selectively rendering histone genes transcribable. These results suggest a possible role for nonhistone chromosomal proteins in mediating the interactions of histones with DNA to render histone genes transcribable.

Base Sequence↗

Human growth hormone DNA sequence and mRNA structure: possible alternative splicing.

We have determined the complete sequence of the human growth hormone (hGH) gene and the position of the mature 5' end of the hGH mRNA within the sequence. Comparison of this sequence with that of a cloned hGH cDNA shows that the gene is interrupted by four intervening sequences. S1 mapping shows that one of these intervening sequences has two different 3' splice sites. These alternate splicing pathways generate hGH peptides of different sizes which are found in normal pituitaries. Comparison of sequences near the 5' end of the hGH mRNA with a similar region of the alpha subunit of the human glycoprotein hormones reveals an unexpected region of homology between these otherwise unrelated peptide hormones.

Amino Acid Sequence↗

Control of expression of the herpes simplex virus thymidine kinase gene in biochemically transformed cells.

A series of cell lines was constructed by transformation of murine LTK- cells with a family of deletion mutants of the herpes simplex virus (HSV) thymidine kinase (tk) gene. These mutants, differing in the extent of 5' sequence flanking the coding region for tk, varied in the frequency with which they were able to convert tk- cells to the tk+ phenotype. Converted cell lines were analysed for tk DNA sequences, tk mRNA sequences, the 5' terminus of tk-specific transcripts and for their ability to respond to a signal provided in trans by infecting tk- virus (transactivation). The results of these analyses reveal that transformation efficiency correlates inversely with the extent of 5' flanking information. Thus mutants retaining less than 109 bp of 5' sequences transform less efficiently than those that retain at least 109 bp. Cell lines established by transformation with mutants retaining the proximal 109 bp contain relatively few copies of tk DNA whereas those which arose as a result of transformation with mutant DNA containing less than 109 bp generally contained multiple copies of tk DNA. Analyses of tk-specific transcripts revealed that cell lines derived from plasmids that transformed efficiently synthesized an mRNA which was indistinguishable by its size or 5' end from infected cell mRNA. Cell lines established by plasmids that were inefficient at transformation accumulated truncated mRNAs that initiated at aberrant start sites. The presence of the 5' 109 bp block was required for transformants to increase the level of tk mRNA and enzyme when infected with a tk- deletion mutant of HSV. We also show that transactivation does not alter the initiation site of the tk mRNA synthesized by transformants.

Animals↗

Methacarn fixation: a novel tool for analysis of gene expressions in paraffin-embedded tissue specimens.

To establish a quantitative method for analysis of gene expressions in small areas of tissue after paraffin embedding, preliminary validation experiments with RT-PCR and Western blotting were performed using methacarn-fixed rodent tissues and a cultured PC12 cell line. A total RNA yield of 52 +/- 15 ng/mm2, sufficient for a quantitative RT-PCR of many genes, could be extracted from a deparaffinized 10-microm-thick rat-liver section by a simple, single-step extraction method. The low concentration of contaminating genomic DNA and the resolution of ribosomal RNAs in RNA gel proved the purity and integrity of the extracted RNA samples, allowing PCR amplification of a long mRNA sequence and mRNA species expressing low copy numbers. PCR amplification of mRNA-derived target gene fragments could be achieved by optimizing the amount of total RNA for reverse transcription and the number of subsequent PCR cycles for each gene. By this validation, organ- and sex-specific mRNA expression could be detected in methacarn-fixed paraffin-embedded tissues without additional DNase treatment of RNA samples. RT-PCR analysis could also be performed with total RNA extracted from deparaffinized tissue dissected with a laser capture microdissection system. In addition, extraction of protein yielded 4.9 +/- 2.1 microg/mm2 from a 10-microm-thick rat-liver section, allowing a quantitative expression analysis of protein by Western blotting. Thus, in addition to its advantages for immunohistochemistry, methacarn-fixed paraffin-embedded tissue has benefits for analysis of both RNAs and proteins in the cells of histologically defined areas.

Acetic Acid↗

Temporal multiomics gene expression data of human embryonic stem cell-derived cardiomyocyte differentiation.

Human embryonic stem cells (hESCs) serve as a valuable in vitro model for studying early human developmental processes due to their ability to differentiate into all three germ layers. Here, we present a comprehensive multi-omics dataset generated by differentiating hESCs into cardiomyocytes via the mesodermal lineage, collecting samples at 10 distinct time points. We measured mRNA levels by mRNA sequencing (mRNA-seq), translation levels by ribosome profiling (Ribo-seq), and protein levels by quantitative mass spectrometry-based proteomics. Technical validation confirmed high quality and reproducibility across all datasets, with strong correlations between replicates. This extensive dataset provides critical insights into the complex regulatory mechanisms of cardiomyocyte differentiation and serves as a valuable resource for the research community, aiding in the exploration of mammalian development and gene regulation.

Humans↗

Piliation control mechanisms in Neisseria gonorrhoeae.

Gonococci (Gc) undergo pilus+ to pilus- "phase transitions" readily in vitro. In the present study we sequenced pilin mRNA from reverting, pilus- Gc by oligonucleotide primer extension and compared these pilin mRNA sequences with those expressed by their pilus+ predecessors and pilus+ revertants. The results suggest that genetic rearrangement within the pilin structural gene can generate defective pilin gene products, resulting in a pilus- phenotype. These pilus- Gc give rise to pilus+ revertants upon reconstitution of their modified pilin gene.

Bacterial Outer Membrane Proteins↗

Human immunodeficiency virus 1 tat protein binds trans-activation-responsive region (TAR) RNA in vitro.

tat, the trans-activator protein for human immunodeficiency virus 1 (HIV-1), has been expressed in Escherichia coli from synthetic genes. Purified tat binds specifically to HIV-1 trans-activation-responsive region (TAR) RNA in gel-retardation, filter-binding, and immunoprecipitation assays. tat does not bind detectably to antisense TAR RNA sequences, cellular mRNA sequences, variant TAR RNA sequences with altered stem-loop structures, or TAR DNA.

Base Sequence↗

Three-dimensional laser-scanning confocal microscopy of in situ hybridization in the skin.

In situ hybridization is an important tool in molecular and developmental biology to detect specific nucleic acid sequences (either mRNA or DNA) within cells. This technique is especially applicable to tissue sections since it provides information about the spatial distribution of DNA or mRNA sequences. However, previous studies utilizing in situ hybridization in the skin were hampered by a high degree of nonspecific background, which has made interpretation of the results difficult. In this paper, we demonstrate how refinements in in situ hybridization techniques, combined with laser-scanning confocal microscopy, significantly reduce nonspecific background and produce improved resolution of in situ hybridization in skin specimens. The sensitive detection method of laser-scanning confocal microscopy allows three-dimensional localization of S35 radioactive-labeled riboprobes within the emulsion of specimens, which is not possible with conventional bright or dark field light microscopy.

Dendritic Cells↗

Improved Method for Recovery of mRNA from Aquatic Samples and Its Application to Detection of mer Expression.

Previously described methods for extraction of mRNA from environmental samples may preclude detecting transcripts from genes that were present in low abundance in aquatic bacterial communities. By combining a boiling sodium dodecyl sulfate-diethylpyrocarbonate lysis step with acid-guanidinium extraction, we improved recovery of target mRNA from both pure cultures and environmental samples. The most significant advantage of the new protocol is that it is easily adapted to yield high recovery of mRNA from 142-mm-diameter flat filters and high-capacity cartridge filters. The lysis and extraction procedures are more rapid than previously described methods, and many samples can be handled at once. RNA extracts have been shown to be free of contaminating DNA. The lysis procedure does not damage target mRNA sequences, and mRNA can be detected from fewer than 10 bacterial cells. We used the new method to examine transcripts of genes responsible for detoxification of mercurial compounds. Induction of merA (specifying mercuric reductase) transcripts in stationary-phase Pseudomonas aeruginosa containing Tn501 occurred within 60 s of HgCl(2) addition and was proportional to the amount of Hg(II) added. The new technique also allowed the detection of merA transcripts from the microbial community of a mercury-contaminated pond (Reality Lake, Oak Ridge, Tenn.). Significant differences in merA transcript abundance were observed between different locations associated with the lake. The results indicate that the new method is simple and rapid and can be applied to the study of mer gene expression of aquatic communities in their natural habitats.

Journal Article↗

A first-generation map of the turkey genome.

A primary linkage map of the domestic turkey (Meleagris gallopavo) was developed by segregation analysis of genetic markers within a backcross family. This reference family includes 84 offspring from one F1 sire mated to two dams. Genomic DNA was digested using one of five restriction enzymes, and restriction fragment length polymorphisms were detected on Southern blots using probes prepared from 135 random clones isolated from a whole-embryo cDNA library. DNA sequence was subsequently determined for 114 of these cDNA clones. Sequence comparisons were done using BLAST searches of the GenBank database, and redundant sequences were eliminated. High similarity was found between 23% of the turkey sequences and mRNA sequences reported for the chicken. The current map, based on expressed genes, includes 138 loci, encompassing 113 loci arranged into 22 linkage groups and an additional 25 loci that remain unlinked. The average distance between linked markers is 6 cM and the longest linkage group (17 loci) measures 131 cM. The total map distance contained within linkage groups is 651 cM. The present map provides an important framework for future genome mapping in the turkey.

Animals↗

Structure of apolipoprotein B-100 of human low density lipoproteins.

We have analyzed low density lipoproteins (LDL) apolipoprotein (apop) B structure by direct sequence analysis of LDL apo B-100 tryptic peptides. Native LDL were digested with trypsin, and the products were fractionated on a Sephadex G-50 column. The partially digested apo B-100 still associated with lipids was recovered in the void volume (designated trypsin-nonreleasable, TN, peptides). The released peptides (designated trypsin-releasable, TR, peptides) in subsequent peaks were repurified on two successive high-performance liquid chromatography (HPLC) columns. The TN peak was delipidated and redigested with trypsin, and the resulting peptides were purified on two successive HPLC columns. Using this approach, we sequenced over 88% of LDL apo B-100, extending and refining our previous study (Nature 1986;323:738-742) which covered 52% of the protein. TN peptides made up 31%, and the TR peptides, 34% of the apo B-100 sequence; 23.7% were found under both TN and TR categories. Based on its differential trypsin releasability, apo B-100 can be divided into five domains: 1) residues 1----1000, largely TR; 2) residues 1001----1700, alternating TR and TN; 3) residues 1701----3070, largely TN; 4) residues 3071----4100, mainly TR and mixed; and 5) residues 4101----4536, almost exclusively TN. Domain 1 contained 14 of the 25 Cys residues in apo B. Domain 4 encompassed seven N-glycosylation sites, and contained the putative receptor binding domains. All 19 potential N-glycosylation sites were directly sequenced: 16 were found to be glycosylated and three were not. Three pairs of disulfide bridges were also mapped. Finally, a combination of cDNA sequencing, direct mRNA sequencing, and comparison of published apo B-100 sequences allowed us to identify specific amino acid residues within apo B-100 that seem to represent bona fide allelic variations. Our study provides information on LDL apo B-100 structure that will be important to our understanding of its conformation and metabolism.

Amino Acid Sequence↗

Detection of minimal residual disease in acute leukemia by immunological marker analysis and polymerase chain reaction.

Detection of minimal residual disease (MRD) can be useful for adaptation or stratification of treatment in acute leukemia patients and may finally result in individualization of treatment protocols. Although leukemic cells generally have immunophenotypes comparable to their normal counterparts, it is possible to use immunological marker analysis for the detection of MRD based on the assumption that the presence of positive cells outside their normal breeding sites and 'homing areas' is indicative of malignancy. This approach can be used for the detection of MRD in blood and bone marrow of patients with a terminal deoxynucleotidyl transferase (TdT) positive T-cell acute lymphoblastic leukemia (ALL) and patients with a TdT+ acute myeloid leukemia (AML) as well as in cerebrospinal fluid of patients with a TdT+ leukemia. In other types of acute leukemias, immunological marker analysis generally does not allow detection of low frequencies of malignant cells, but in a part of them the polymerase chain reaction (PCR) technique may be valuable. The PCR technique allows the amplification of tumor-specific DNA sequences or mRNA sequences (after reverse transcription into cDNA), if the flanking sequences are well-defined. This PCR-mediated amplification can detect specific sequences which are derived from only a few malignant cells between many normal cells. Well-defined chromosome translocations have been used as tumor-specific markers, such as t(9;22). An advantage of using specific chromosome aberrations as tumor-specific markers is their stability during the disease course. However, only 10-15% of ALL and 25-30% of AML have a specific chromosome translocation and in a large part of them the precise breakpoints are not (yet) known. Recent studies indicate that it is possible to detect MRD in acute leukemias by use of PCR-mediated amplification of the junctional regions of rearranged immunoglobulin (Ig) and T-cell receptor (TcR) genes, using variable (V) and joining (J) gene-specific oligonucleotides as primers. Major pitfalls of this application are the occurrence of multiple rearrangements at diagnosis (oligoclonality) and changes in rearrangement patterns at relapse (clonal evolution), which will lead to false negative results of this MRD-PCR technique. In conclusion, the technique of choice for the detection of MRD is dependent on the immunophenotype of the leukemia, the presence of a well-defined chromosome translocation and the presence of a rearranged Ig and/or TcR gene as well as the chance of immunophenotypic shifts and changes in Ig and TcR gene rearrangement patterns.(ABSTRACT TRUNCATED AT 400 WORDS)

Acute Disease↗