Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcript abundance”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Statistical analysis of high-density oligonucleotide arrays: a multiplicative noise model.

MOTIVATION: High-density oligonucleotide arrays (GeneChip, Affymetrix, Santa Clara, CA) have become a standard research tool in many areas of biomedical research. They quantitatively monitor the expression of thousands of genes simultaneously by measuring fluorescence from gene-specific targets or probes. The relationship between signal intensities and transcript abundance as well as normalization issues have been the focus of much recent attention (Hill et al., 2001; Chudin et al., 2002; Naef et al., 2002a). It is desirable that a researcher has the best possible analytical tools to make the most of the information that this powerful technology has to offer. At present there are three analytical methods available: the newly released Affymetrix Microarray Suite 5.0 (AMS) software that accompanies the GeneChip product, the method of Li and Wong (LW; Li and Wong, 2001), and the method of Naef et al. (FN; Naef et al., 2001). The AMS method is tailored for analysis of a single microarray, and can therefore be used with any experimental design. The LW method on the other hand depends on a large number of microarrays in an experiment and cannot be used for an isolated microarray, and the FN method is particular to paired microarrays, such as resulting from an experiment in which each 'treatment' sample has a corresponding 'control' sample. Our focus is on analysis of experiments in which there is a series of samples. In this case only the AMS, LW, and the method described in this paper can be used. The present method is model-based, like the LW method, but assumes multiplicative not additive noise, and employs elimination of statistically significant outliers for improved results. Unlike LW and AMS, we do not assume probe-specific background (measured by the so-called mismatch probes). Rather, we assume uniform background, whose level is estimated using both the mismatch and perfect match probe intensities. RESULTS: We present a new method for GeneChip analysis, based on a statistical model with multiplicative noise. We demonstrated that this method yields results superior to those obtained by the Affymetrix Microarray Suite 5.0 software and to those obtained by the model-based method of Li and Wong (Li and Wong, 2001). The present method eliminates the hard-to-interpret negative expression indices, and the binary 'presence' calls (present or absent) are replaced by the statistical significance (p-value) of gene expression. We have found that thresholding the p-values at the (0.1)(16)-level produces about the same number of 'present' calls as the AMS software. By testing our method on a pair of replicate GeneChips (hybridized with the same cRNA), we found that 95.6% of data points lie within the 1.25-fold interval. In other words, our method had a 4.4% type I error rate at the 1.25-fold level. The error rate of the LW method was 15%, and that of the AMS method was 29%. There were no points outside the 2-fold interval with the present method. Analysis of variance (ANOVA) of another experiment with multiple replicates shows that this reduction of variance is not accompanied by a corresponding reduction of signal. On the contrary, the signal-to-noise ratio (as measured by the distribution of F-statistics) of the present method is on average 3.4-times better than that of AMS, and 1.4-times better than that of Li and Wong.

Algorithms↗

Striping artifact removal in VisiumHD data through nuclear counts modeling.

MOTIVATION: 10x Genomics VisiumHD enables spatial transcriptomics at 2 µm × 2 µm resolution but exhibits slide-specific, non-periodic striping artifacts due to lane-width variability. These multiplicative row/column effects distort bin total counts and can bias downstream analyses. The state-of-the-art destriping approach is the normalization procedure used as a preprocessing step in bin2cell; it applies sequential high-quantile row- then column-wise normalization, which is asymmetric and can introduce edge effects/macro-stripes and distortions of large-scale total-count structure. RESULTS: We propose a statistical destriping approach that leverages nuclei segmentation from the co-registered H&E image. Assuming transcript abundance is constant within each nucleus, we model bin counts with a negative binomial distribution whose mean is a product of a nucleus-specific concentration and row- and column-specific stripe-factors reflecting lane-width variation. We fit all parameters in a generalized linear modeling framework with cross-validated regularization on stripe-factors and iterative dispersion estimation, and use the fitted parameters to correct the observed counts into a destriped image. On synthetic data with known ground truth, our method improves stripe-factor estimation accuracy and reduces error in corrected counts relative to bin2cell and bin2cell-derived baselines. Across four public VisiumHD slides, it consistently lowers striping intensity while substantially better preserving biological signal present in the large-scale global count structure and avoiding the artifacts introduced by other methods. AVAILABILITY AND IMPLEMENTATION: All source code and links to publicly available data used for this study are available at https://github.com/paolamalsot/destriping-GLM.

Artifacts↗

Using a calibration experiment to assess gene-specific information: full Bayesian and empirical Bayesian models for two-channel microarray data.

MOTIVATION: Microarray studies permit to quantify expression levels on a global scale by measuring transcript abundance of thousands of genes simultaneously. A difficulty when analysing expression measures is how to model variability for the whole set of genes. It is usually unrealistic to assume a common variance for each gene. Several approaches to model gene-specific variances are proposed. We take advantage of calibration experiments, in which the probes hybridized on the two channels come from the same population (self-self experiment). In this case it is possible to estimate the gene-specific variance, to be incorporated in comparative experiments on the same tissue, cellular line or species. RESULTS: We present two approaches to introduce prior information on gene-specific variability from a calibration experiment: an empirical Bayes model and a full Bayesian hierarchical model. We apply the methods in the analysis of human lipopolysaccharide-stimulated leukocyte experiments. AVAILABILITY: The calculations are implemented in WinBugs. The codes are available on request from the authors.

Algorithms↗

The psbO gene for 33-kDa precursor polypeptide of the oxygen-evolving complex in Arabidopsis thaliana--nucleotide sequence and control of its expression.

The 33-kDa polypeptide of the oxygen-evolving complex of photosystem II is nuclear-encoded. The single psbO gene of Arabidopsis thaliana, as suggested by Southern hybridization, has been isolated from the genomic library and sequenced. The sequence analysis has revealed that the psbO gene harbors two introns and encodes a precursor polypeptide of 332 amino acid residues; the first 85 amino acid residues represent the transit peptide and the following 247 amino acids constitute the mature polypeptide. The hydrophilic nature of the 33-kDa protein is confirmed by the presence of 27% charged residues. Northern analysis of the total RNA from Arabidopsis indicates that a 1.2-kb transcript represents the psbO gene. It is expressed in a tissue-specific manner -- the steady-state transcript levels being highest in the leaves and virtually undetectable in the roots. Also, expression of the psbO gene is development-dependent and regulated by light in young Arabidopsis seedlings. In a constitutively photomorphogenic mutant of Arabidopsis, pho2 (plumular hook open 2), the psbO gene is de-repressed in young, dark-grown seedlings, resulting in increased transcript abundance compared to the wild-type. These studies, thus, define the influence of at least one regulatory component for psbO expression.

Amino Acid Sequence↗

Expressed sequence tags from immature female sexual organ of a liverwort, Marchantia polymorpha.

A total of 970 expressed sequence tag (EST) clones were generated from immature female sexual organ of a liverwort, Marchantia polymorpha. The 376 ESTs resulted in 123 redundant groups, thus the total number of unique sequences in the EST set was 717. Database search by BLAST algorithm showed that 302 of the unique sequences shared significant similarities to known nucleotide or amino acid sequences. Six unique sequences showed significant similarities to genes that are involved in flower development and sexual reproduction, such as cynarase, fimbriata-associated protein and S-receptor kinase genes. The remaining unique 415 sequences have no significant similarity with any database-registered genes or proteins. The redundant 123 ESTs implied the presence of gene families and abundant transcripts of unknown identity. Analyses of the coding sequences of 61 unique sequences, which contained no ambiguous bases in the predicted coding regions, highly homologous to known sequences at the amino acid level with a similarity score greater than 400, and with stop codons at similar positions as their possible orthologues, indicated the presence of biased codon usage and higher GC content within the coding sequences (50.4%) than that within 3' flanking sequences (41.9%).

Amino Acid Sequence↗

Natural synthesis of a DNA-binding protein from the C-terminal domain of DNA gyrase A in Borrelia burgdorferi.

We have identified a 34 kDa DNA-binding protein with an HU-like activity in the Lyme disease spirochete Borrelia burgdorferi. The 34 kDa protein is translated from an abundant transcript initiated within the gene encoding the A subunit of DNA gyrase. Translation of the 34 kDa protein starts at residue 499 of GyrA and proceeds in the same reading frame as full-length GyrA, resulting in an N-terminal-truncated protein. The 34 kDa GyrA C-terminal domain, although not homologous, substitutes for HU in the formation of the Type 1 complex in Mu transposition, and complements an HU-deficient strain of Escherichia coli. This is the first example of constitutive expression of two gene products in the same open reading frame from a single gene in a prokaryotic cellular system.

Amino Acid Sequence↗

Translational selection and yeast proteome evolution.

The primary structures of peptides may be adapted for efficient synthesis as well as proper function. Here, the Saccharomyces cerevisiae genome sequence, DNA microarray expression data, tRNA gene numbers, and functional categorizations of proteins are employed to determine whether the amino acid composition of peptides reflects natural selection to optimize the speed and accuracy of translation. Strong relationships between synonymous codon usage bias and estimates of transcript abundance suggest that DNA array data serve as adequate predictors of translation rates. Amino acid usage also shows striking relationships with expression levels. Stronger correlations between tRNA concentrations and amino acid abundances among highly expressed proteins than among less abundant proteins support adaptation of both tRNA abundances and amino acid usage to enhance the speed and accuracy of protein synthesis. Natural selection for efficient synthesis appears to also favor shorter proteins as a function of their expression levels. Comparisons restricted to proteins within functional classes are employed to control for differences in amino acid composition and protein size that reflect differences in the functional requirements of proteins expressed at different levels.

Adaptation, Physiological↗

Gene expression is stable despite widespread cis and trans regulatory divergence in Saccharomyces yeasts.

Regulatory evolution can alter phenotypes, but cis- and trans-regulatory mechanisms may also diverge extensively while total transcript abundance remains stable. Comparisons of parental expression with allele-specific expression in F1 hybrids provide a framework for separating cis- and trans-regulatory effects because both parental alleles are measured in a shared trans-regulatory environment. Here, we analyzed RNA sequencing data from Saccharomyces cerevisiae, Saccharomyces paradoxus, and their F1 hybrid. Among the 4,164 genes with sufficient allele-specific support for strict classification, 2,134 (51.2%) showed detectable cis and/or trans regulatory divergence. However, hybrid expression remained largely conserved, with 81.5% of genes not significantly different from either parent. Compensatory cis-trans divergence predominated over reinforcing divergence; cross-replicate estimation reduced the apparent magnitude of this excess, but opposite-sign effects remained predominant in all 20 non-overlapping replicate comparisons. To connect gene expression to genome sequence, we analyzed the strongly cis-diverged locus LYS2 and found species differences in promoter architecture, including an S. cerevisiae-specific AT-rich insertion, altered spacing among candidate regulatory features, and a promoter-proximal TATA-like element unique to S. cerevisiae. Sequence-based nucleosome prediction suggests that these differences create a broader promoter-proximal nucleosome-depleted region in S. cerevisiae than in S. paradoxus. We also quantified allele-resolved intron retention and found that allele-resolved intron retention was broadly conserved, with only rare locus-specific hybrid-associated shifts. Together, these results show that regulatory divergence is widespread but often buffered in the hybrid, whereas intron-retention divergence is comparatively limited.

Saccharomyces↗

Gene expression profiling in human preadipocytes and adipocytes by microarray analysis.

Uncontrolled expansion of adipose tissue leads to obesity, a public health epidemic affecting >30% of adult Americans. Adipose mass increases in part through the recruitment and differentiation of an existing pool of preadipocytes (PA) into adipocytes (AD). Most studies investigating adipogenesis used primarily murine cell lines; much less is known about the relevant processes that occur in humans. Therefore, characterization of genes associated with adipocyte development is key to understanding the pathogenesis of obesity and developing treatments for this disorder. To address this issue, we performed large-scale analyses of human adipose gene expression using microarray technology. Differential gene expression between PA and AD was analyzed in 6 female patients using human cDNA microarray slides and data analyzed using the Stanford Microarray Database. Statistical analysis for the gene expression was performed using the SAS mixed models. Compared with PA, several genes involved in lipid metabolism were overexpressed in AD, including fatty acid binding protein, adipose differentiation-related protein, lipoprotein lipase, perilipin, and adipose most abundant transcript 1. Novel genes expressed in adipocytes included E2F5 transcriptional factor and SMARC (SWI/SNF-related, matrix associated, actin-dependent regulator of chromatin). PA predominantly expressed genes encoding extracellular matrix components such as fibronectin, matrix metalloprotein, and novel proteins such as lysyl oxidase. Despite the high differential expression of some of these genes, many did not differ significantly likely due to high variability and limited statistical power. A comprehensive list of differential gene expression is presented according to cellular function. In conclusion, these studies offer an overview of the gene expression profiles in PA and AD and identify new genes with potentially important functions in adipose tissue development and obesity that merit further investigation.

Adipocytes↗

Systemic signalling of environmental cues in Arabidopsis leaves.

Light intensity and atmospheric CO2 partial pressure are two environmental signals known to regulate stomatal numbers. It has previously been shown that if a mature Arabidopsis leaf is supplied with either elevated CO2 (750 ppm instead of ambient at 370 ppm) or reduced light levels (50 micromol m-2 s-1 instead of 250 micromol m-2 s-1), the young, developing leaves that are not receiving the treatment grow with a stomatal density as if they were exposed to the treatment. But the signal(s) that it is believed is generated in the mature leaves and transmitted to developing leaves are largely unknown. Photosynthetic rates of treated, mature Arabidopsis leaves increased in elevated CO2 and decreased when shaded, as would be expected. Similarly, the levels of sugars (glucose, fructose, and sucrose) in the treated mature leaves increased in elevated CO2 and decreased with shade treatment. The levels of sugar in developing leaves were also measured and it was found that they mirrored this result even though they were not receiving the shade or elevated CO2 treatment. To investigate the effect of these treatments on global gene expression patterns, transcriptomics analysis was carried out using Affymetrix, 22K, and ATH1 arrays. Total RNA was extracted from the developing leaves after the mature leaves had received either the ambient control treatment, the elevated CO2 treatment, or the shade treatment, or both elevated CO2 and shade treatments for 2, 4, 12, 24, 48, or 96 h. The experiment was replicated four times. Two other experiments were also conducted, one to compare and contrast gene expression in response to plants grown at elevated CO2 and the other to look at the effect of these treatments on the mature leaf. The data were analysed and 915 genes from the untreated, signalled leaves were identified as having expression levels affected by the shade treatment. These genes were then compared with those whose transcript abundance was affected by the shade treatment in the mature treated leaves (1181 genes) and with 220 putative 'stomatal signalling' genes previously identified from studies of the yoda mutant. The results of these experiments and how they relate to environmental signalling are discussed, as well as possible mechanisms for systemic signalling.

Acclimatization↗

Two new cysteine proteinases with specific expression patterns in mature and senescent tobacco (Nicotiana tabacum L.) leaves.

Cysteine proteinases are involved in various physiological and developmental processes in plants. Two cDNAs from senescent and non-senescent tobacco leaves were isolated with degenerate primers designed from conserved regions of plant senescence-associated cysteine proteinases using rapid amplification of cDNA ends (RACE). Both sequences encode papain-like cysteine proteinases: the 833 bp fragment (NtCP1) encoding a C-terminus partial sequence of a putative tobacco cysteine proteinase gene whereas the 1300 bp fragment (NtCP2) is a full-length cysteine proteinase. On the amino acid sequence level, NtCP1 has a high similarity with other senescence-associated cysteine proteinases. It is expressed only in senescent leaves. It is not induced in mature green leaves upon exposure to drought or heat. These results suggest that it might be a good developmental senescence marker in tobacco. By contrast, NtCP2 has a high similarity to KDEL-tailed cysteine proteinases and is expressed in mature green leaves. Both drought and heat decreased NtCP2 transcript abundance in mature green leaves. It is concluded that NtCP1 is a senescence-specific cysteine proteinase whereas NtCP2 fulfils roles in green leaves that might be similar to those of KDEL-tailed cysteine proteinases involved, for example, in programmed cell death.

Amino Acid Sequence↗

Common pattern of evolution of gene expression level and protein sequence in Drosophila.

Sequence divergence scaled by variation within species has been used to infer the action of selection upon individual genes. Applying this approach to expression, we compared whole-genome whole-body RNA levels in 10 heterozygous Drosophila simulans genotypes and a pooled sample of 10 D. melanogaster lines using Affymetrix Genechip. For 972 genes expressed in D. melanogaster, the transcript level was below detection threshold in D. simulans, which may be explained either by sequence divergence between the primers on the chip and the mRNA transcripts or by down-regulation of these genes. Out of 6,707 genes that were expressed in both species, transcript level was significantly different between species for 534 genes (at P < 0.001). Genes whose expression is under stabilizing selection should exhibit reduced genetic variation within species and reduced divergence between species. Expression of genes under directional selection in D. simulans should be highly divergent from D. melanogaster, while showing low genetic variation in D. simulans. Finally, the genes with large variation within species but modest divergence between species are candidates for balancing selection. Rapidly diverging, low-polymorphism genes included those involved in reproduction (e.g., Mst 3Ba, 98Cb; Acps 26Aa, 63F; and sperm-specific dynein). Genes with high variation in transcript abundance within species included metallothionein and hairless, both hypothesized to be segregating in nature because of gene-by-environment interactions. Further, we compared expression divergence and DNA substitution rate in 195 genes. Synonymous substitution rate and expression divergences were uncorrelated, whereas there was a significant positive correlation between nonsynonymous substitution rate and expression divergence. We hypothesize that as a substantial fraction of nonsynonymous divergence has been shown to be adaptive, much of the observed expression divergence is likewise adaptive.

Amino Acid Sequence↗

Interleukin-11 inhibits expression of insulin-like growth factor binding protein-5 mRNA in decidualizing human endometrial stromal cells.

Differentiation of endometrial stromal cells into decidual cells is essential for successful embryo implantation. Interleukin (IL)-11 signalling is critical for normal decidualization in the mouse. The expression of IL-11 and its receptors during the menstrual cycle, and the effect of exogenous IL-11 on the decidualization of human endometrial stromal cells in vitro, suggests a role for this cytokine in human decidualization. As the downstream target genes of IL-11 are also likely to be critical mediators of this process, this study aimed to identify genes regulated by IL-11 in decidualizing human endometrial stromal cells in vitro. Stromal cells isolated from endometrial biopsies were decidualized with 17beta estradiol (E) and medroxyprogesterone acetate (EP) in the presence or absence of exogenous IL-11, and total RNA used for cDNA microarray analysis and real-time RT-PCR. Microarray analysis revealed 16 up-regulated and 11 down-regulated cDNAs in EP + IL-11-treated compared with EP-treated cells. The most down-regulated gene was insulin-like growth factor binding protein-5 (IGFBP-5) (3.6-fold). Using real-time RT-PCR, IL-11 was confirmed to decrease IGFBP-5 transcript abundance 102-fold (P = 0.016; n = 6). No difference in IGFBP-5 immunostaining intensity was detected in stromal cells decidualized in the presence or absence of IL-11, and there was no effect of exogenous IGFBP-5 on the progression of steroid-induced in vitro decidualization. Interactions between IL-11 and its target genes, including IGFBP-5, may contribute to the regulation of decidualization and/or mediate communication between the decidua and invading trophoblast at implantation.

Cell Differentiation↗

The nucleotide sequence of a segment of Trypanosoma brucei mitochondrial maxi-circle DNA that contains the gene for apocytochrome b and some unusual unassigned reading frames.

The nucleotide sequence of a 2.5-kb segment of the maxi-circle of Trypanosoma brucei mtDNA has been determined. The segment contains the gene for apocytochrome b, which displays about 25% homology at the amino acid level to the apocytochrome b gene from fungal and mammalian mtDNAs. Northern blot and S1 nuclease analyses have yielded accurate map positions of an RNA species in an area that coincides with the reading frame. The segment also contains two pairs of overlapping unassigned reading frames, which lack homology with any known mitochondrial gene or URF. The DNA sequence in these areas is AG-rich (70%), resulting in URFs with an unusually high level of glycine and charged amino acids (60%). They may not encode proteins, in spite of their size and the fact that abundant transcripts are mapped in these areas.

Amino Acid Sequence↗

Characterization of an unique RNA initiated immediately upstream from human alpha 1 globin gene in vivo and in vitro: polymerase II-dependence, tissue specificity, and subcellular location.

We have identified an abundant transcript initiated upstream from the canonical cap site of human alpha 1 globin gene in bone marrow cells and in COS-7 cells transfected with an alpha 1 globin gene-containing plasmid. Similar to the major alpha 1 globin transcript, this upstream RNA is present almost exclusively in the cytoplasm of the transfected COS-7 cells. It is also synthesized efficiently in vitro by RNA polymerase II in the nuclear extracts prepared from a Hela cell line and an erythroleukemia cell line, K562. RNAs isolated from these cell lines, however, do not contain this upstream transcript. The putative 5' end of the alpha 1 globin upstream RNA is mapped by primer extension to base -45, which is located in between the CCAAT and TATA boxes. The synthesis of this RNA in vitro and in vivo, and the close proximity of its 5' end to the promoter of the alpha 1 globin gene suggest a common mechanism regulating the transcriptional initiation of both the upstream and the major alpha 1 globin RNAs.

Animals↗

Characterization of yeast mitochondrial RNase P: an intact RNA subunit is not essential for activity in vitro.

We have previously described a mitochondrial activity that removes 5' leaders from yeast mitochondrial precursor tRNAs and suggested that it is a mitochondrial RNase P. Here we demonstrate that the cleavage reaction results in a 5' phosphate on the tRNA product and thus the activity is analogous to that of other RNase Ps. A mitochondrial gene called the tRNA synthesis locus encodes an A + U-rich RNA required for this activity in vivo. Two regions of this RNA display sequence similarity to conserved sequences in bacterial RNase P RNAs. This sequence similarity coupled with the analogous activities of the enzymes has led us to conclude that the RNAs are homologous and that the tRNA synthesis locus does code for the mitochondrial RNase P RNA subunit. The smallest and most abundant transcript of the tRNA synthesis locus is 490 nucleotides long. However, during purification of the holoenzyme, RNA is degraded and pieces of the original RNA are sufficient to support RNase P activity in vitro.

Bacillus subtilis↗

Unusual features of the retroid element PAT from the nematode Panagrellus redivivus.

The PAT retroid transposable elements differ from other retroids in that they have a 'split direct repeat' structure, i.e., and internal 300bp sequence is found repeated, about one half at each element extremity. A very abundant transcript of about 900 nt, the start of which maps to the preferentially deleted portion of PAT elements, is detected on total Panagrellus redivius RNA bearing Northern blots. A potentially corresponding ORF encodes a protein of 265 residues having a carboxy terminal Cystein motif, believed to be exclusively characteristic of the GAG protein in retoid elements. A much fainter, 1800nt long transcript, is also detected on Northern blots and maps slightly downstream of the first ORF. The predicted protein sequence of this region bears motifs typical of reverse transcriptase and RNaseH, as found in the Pol genes of retroid elements. Peptide motif similarities are greatest with the DIRS-1 element derived from Dictyostelium discoideum. The possibility of using PAT elements as transposon tagging system for Caenorhabditis elegans is discussed.

Amino Acid Sequence↗

Cloning, expression and localization of an RNA helicase gene from a human lymphoid cell line with chromosomal breakpoint 11q23.3.

A gene encoding a putative human RNA helicase, p54, has been cloned and mapped to the band q23.3 of chromosome 11. The predicted amino acid sequence shares a striking homology (75% identical) with the female germline-specific RNA helicase ME31B gene of Drosophila. Unlike ME31B, however, the new gene expresses an abundant transcript in a large number of adult tissues and its 5' non-coding region was found split in a t(11;14)(q23.3;q32.3) cell line from a diffuse large B-cell lymphoma.

Amino Acid Sequence↗