Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genes, Overlapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Heterogeneity within animal thioredoxin reductases. Evidence for alternative first exon splicing.

Animal thioredoxin reductases (TRs) are selenocysteine-containing flavoenzymes that utilize NADPH for reduction of thioredoxins and other protein and nonprotein substrates. Three types of mammalian TRs are known, with TR1 being a cytosolic enzyme, and TR3, a mitochondrial enzyme. Previously characterized TR1 and TR3 occurred as homodimers of 55-57-kDa subunits. We report here that TR1 isolated from mouse liver, mouse liver tumor, and a human T-cell line exhibited extensive heterogeneity as detected by electrophoretic, immunoblot, and mass spectrometry analyses. In particular, a 67-kDa band of TR1 was detected. Furthermore, a novel form of mouse TR1 cDNA encoding a 67-kDa selenoprotein subunit with an additional N-terminal sequence was identified. Subsequent homology analyses revealed three distinct isoforms of mouse and rat TR1 mRNA. These forms differed in 5' sequences that resulted from the alternative use of the first three exons but had common downstream sequences. Similarly, expression of multiple mRNA forms was observed for human TR3 and Drosophila TR. In these genes, alternative first exon splicing resulted in the formation of predicted mitochondrial and cytosolic proteins. In addition, a human TR3 gene overlapped with the gene for catechol-O-methyltransferase (COMT) on a complementary DNA strand, such that mitochondrial TR3 and membrane-bound COMT mRNAs had common first exon sequences; however, transcription start sites for predicted cytosolic TR3 and soluble COMT forms were separated by approximately 30 kilobases. Thus, this study demonstrates a remarkable heterogeneity within TRs, which, at least in part, results from evolutionary conserved genetic mechanisms employing alternative first exon splicing. Multiple transcription start sites within TR genes may be relevant to complex regulation of expression and/or organelle- and cell type-specific location of animal thioredoxin reductases.

Alternative Splicing↗

Characterization of two Azospirillum brasilense Sp7 plasmid genes homologous to Rhizobium meliloti nodPQ.

Bacteria belonging to the Azospirillum genus are nitrogen fixers that colonize the roots of grasses, but do not cause the formation of differentiated structures. Sequences from total DNA of several Azospirillum strains are homologous to restriction fragments containing Rhizobium meliloti nodulation genes. A 10-kilobase (kb) EcoRI fragment from A. brasilense Sp7, sharing homology with a 6.8-kb EcoRI fragment carrying nodGEFH and part of nodP of R. meliloti 41, was cloned in pUC18 to yield pAB502. The nucleotide sequence of a 3.5-kb EcoRI-SmaI fragment of the pAB502 insert revealed 60% homology with R. meliloti nodP and nodQ genes. The nodP gene product shares no homology to any known protein sequence. The Azospirillum nodQ gene product shares homology with a family of initiation and elongation factors as does the R. meliloti nodQ gene product. Since the nodQ gene overlaps the nodP gene, the two genes might be cotranscribed. Azospirillum contains large plasmids, and the nodPQ genes were found on the 90-MDa plasmid (p90). A translational nodP-lacZ fusion was constructed in the broad host range plasmid pGD926. No beta-galactosidase activity was detected in Escherichia coli, but the fusion was functional in Azospirillum and constitutively expressed. Deletions and mutations of nodPQ did not modify growth, nitrogen fixation, or interaction with wheat seedlings.

Amino Acid Sequence↗

Molecular characterization of pssCDE genes of Rhizobium leguminosarum bv. trifolii strain TA1: pssD mutant is affected in exopolysaccharide synthesis and endocytosis of bacteria.

We have identified the three genes pssCDE in Rhizobium leguminosarum bv. trifolii TA1. Even though they were almost identical to earlier identified pssCDE genes of R. leguminosarum, they differed in gene lengths and gene overlaps. The predicted gene products of pssCDE genes shared significant homology to prokaryotic glycosyl transferases involved in exopolysaccharide synthesis. The Tn5 insertion in pssD created the nonmucoid mutant that induced non-nitrogen-fixing nodules. The microscopic analysis of the nodules, induced on Trifolium pratense by the pssD133 mutant, showed abnormally enlarged infection threads densely packed with bacteria, which were released from the infection threads in an unusual way. The symbiosomes were observed very rarely and the nodule remained almost empty. Symbiotic phenotype of the pssD133 suggested a correlation between this mutation and defective endocytosis of bacteria into nodule cells.

Endocytosis↗

Overlapping Lsp-2 gene sequences target expression to both the larval and adult Drosophila fat body.

Larval serum protein-2 gene (Lsp-2) of Drosophila melanogaster encodes one of the major hexameric haemolymph proteins of third-instar larvae and a major component of adult serum. Regulated transcription of Lsp-2 results in high-level, ecdysone-stimulated expression throughout the larval fat body and low-level, spatially restricted expression in the adult fat cells. To localize cis-acting regulatory sequences responsible for the stage- and tissue-specific activity at Lsp-2, the expression of Lsp-2-lacZ fusion genes was studied by P element-mediated germline transformation of Drosophila. A 230 base pair larval enhancer, which includes an ecdysone response element (EcRE), specifically targets gene activity to the larval fat body. Although the adult mode of Lsp-2 expression depends on the larval enhancer, additional negative regulatory elements dictate both tissue-specificity and unique spatial restriction within the adult fat body. Implications of these findings for the identification of fat body-specific gene regulatory units in other insects are discussed.

Animals↗

Sequence analysis of scrA and scrB from Streptococcus sobrinus 6715.

The complete nucleotide sequences of Streptococcus sobrinus 6715 scrA and scrB, which encode sucrose-specific enzyme II of the phosphoenolpyruvate-dependent phosphotransferase system and sucrose-6-phosphate hydrolase, respectively, have been determined. These two genes were transcribed divergently, and the initiation codons of the two open reading frames were 192 bp apart. The transcriptional initiation sites were determined by primer extension analysis, and the putative promoter regions of these two genes overlapped partially. The gene encoding enzyme IIScr, scrA, contained 1,896 nucleotides, and the molecular mass of the predicted protein was 66,529 Da. The hydropathy plot of the predicted amino acid sequence indicated that enzyme IIScr was a relatively hydrophobic protein. The gene encoding sucrose-6-phosphate hydrolase, scrB, contained 1,437 nucleotides. The molecular mass of the predicted protein was 54,501 Da, and the encoded enzyme was hydrophilic. The predicted amino acid sequences of the two open reading frames exhibited approximately 45 and 70% identity with those encoded by scrA and scrB, respectively, from Streptococcus mutans GS5. Homology also was observed between the N-terminal region of the S. sobrinus 6715 enzyme IIScr and other enzyme IIs specific for the glucopyranoside molecule, all of which generate glucopyranoside-6-phosphate during translocation and phosphorylation of the respective substrates. The sequence of the C-terminal domain of the S. sobrinus 6715 enzyme IIScr shared significant homology with enzyme IIIGlc from Escherichia coli and Salmonella typhimurium and with the C-terminal domain of enzyme IIBgl from E. coli, indicating that the two functional domains, enzyme IIScr and enzyme IIIScr, were covalently linked as a single polypeptide in S. sobrinus 6715. The deduced amino acid sequence of the gene product of S. sobrinus scrB shared strong homology with sucrase from Bacillus subtilis, Klebsiella pneumoniae, and Vibrio alginolyticus, suggesting conservation based on the physiological roles of these proteins.

Amino Acid Sequence↗

Identification of the operon for the sorbitol (Glucitol) Phosphoenolpyruvate:Sugar phosphotransferase system in Streptococcus mutans.

Transposon mutagenesis and marker rescue were used to isolate and identify an 8.5-kb contiguous region containing six open reading frames constituting the operon for the sorbitol P-enolpyruvate phosphotransferase transport system (PTS) of Streptococcus mutans LT11. The first gene, srlD, codes for sorbitol-6-phosphate dehydrogenase, followed downstream by srlR, coding for a transcriptional regulator; srlM, coding for a putative activator; and the srlA, srlE, and srlB genes, coding for the EIIC, EIIBC, and EIIA components of the sorbitol PTS, respectively. Among all sorbitol PTS operons characterized to date, the srlD gene is found after the genes coding for the EII components; thus, the location of the gene in S. mutans is unique. The SrlR protein is similar to several transcriptional regulators found in Bacillus spp. that contain PTS regulator domains (J. Stülke, M. Arnaud, G. Rapoport, and I. Martin-Verstraete, Mol. Microbiol. 28:865-874, 1998), and its gene overlaps the srlM gene by 1 bp. The arrangement of these two regulatory genes is unique, having not been reported for other bacteria.

Amino Acid Sequence↗

Use of in vivo expression technology to identify genes important in growth and survival of Pseudomonas fluorescens Pf0-1 in soil: discovery of expressed sequences with novel genetic organization.

Studies were undertaken to determine the genetic needs for the survival of Pseudomonas fluorescens Pf0-1, a gram-negative soil bacterium potentially important for biocontrol and bioremediation, in soil. In vivo expression technology (IVET) identified 22 genes with elevated expression in soil relative to laboratory media. Soil-induced sequences included genes with probable functions of nutrient acquisition and use, and of gene regulation. Ten sequences, lacking similarity to known genes, overlapped divergent known genes, revealing a novel genetic organization at those soil-induced loci. Mutations in three soil-induced genes led to impaired early growth in soil but had no impact on growth in laboratory media. Thus, IVET studies have identified sequences important for soil growth and have revealed a gene organization that was undetected by traditional laboratory approaches.

Bacterial Proteins↗

Nucleotide sequence of AKV murine leukemia virus.

AKV is an endogenous, ecotropic murine leukemia virus that serves as one of the parents of the recombinant; oncogenic mink cell focus-forming viruses that arise in preleukemic AKR mice. I report the 8,374-nucleotide-long sequence of AKV, as determined from the infectious molecular clone AKR-623. The 5'-leader sequence of AKV extends to nucleotide 639, after which lies a long open reading frame encoding the gag and pol gene products. The reading frame is interrupted by a single amber codon separating the gag and pol genes. The pol gene overlaps the env gene within the 3' region of the AKV genome. The nucleotide sequence of the 5' region of AKV reveals the following features. (i) The 5'-leader sequence lacks any AUG codon to initiate translation of gPr80gag, suggesting that gPr80gag is not required for the replication of AKV. (ii) A short portion of the leader region diverges in sequence from the closely related Moloney murine leukemia virus and appears to be related to a sequence highly repeated in eucaryotic genomes. (iii) As in Moloney murine leukemia virus, there is a potential RNA secondary structure flanking the amber codon that separates the gag and pol genes. This structure might function as a regulatory protein binding site that controls the relative levels of synthesis of the gag and pol precursors. The nucleotide sequence of the 3' region of AKV is compared with sequences reported previously from both infectious and noninfectious molecular clones of AKV.

Amino Acid Sequence↗

Analysis of the primary structure of the long terminal repeat and the gag and pol genes of the human spumaretrovirus.

The nucleotide sequence of the human spumaretrovirus (HSRV) genome was determined. The 5' long terminal repeat region was analyzed by strong stop cDNA synthesis and S1 nuclease mapping. The length of the RU5 region was determined and found to be 346 nucleotides long. The 5' long terminal repeat is 1,123 base pairs long and is bound by an 18-base-pair primer-binding site complementary to the 3' end of mammalian lysine-1,2-specific tRNA. Open reading frames for gag and pol genes were identified. Surprisingly, the HSRV gag protein does not contain the cysteine motif of the nucleic acid-binding proteins found in and typical of all other retroviral gag proteins; instead the HSRV gag gene encodes a strongly basic protein reminiscent of those of hepatitis B virus and retrotransposons. The carboxy-terminal part of the HSRV gag gene products encodes a protease domain. The pol gene overlaps the gag gene and is postulated to be synthesized as a gag/pol precursor via translational frameshifting analogous to that of Rous sarcoma virus, with 7 nucleotides immediately upstream of the termination codons of gag conserved between the two viral genomes. The HSRV pol gene is 2,730 nucleotides long, and its deduced protein sequence is readily subdivided into three well-conserved domains, the reverse transcriptase, the RNase H, and the integrase. Although the degree of homology of the HSRV reverse transcriptase domain is highest to that of murine leukemia virus, the HSRV genomic organization is more similar to that of human and simian immunodeficiency viruses. The data justify classifying the spumaretroviruses as a third subfamily of Retroviridae.

Amino Acid Sequence↗

Relationship of the env genes and the endonuclease domain of the pol genes of simian foamy virus type 1 and human foamy virus.

We have molecularly cloned and sequenced a portion of the simian foamy virus type 1 (SFV-1); open reading frames representing the endonuclease domain of the polymerase (pol) and the envelope (env) genes were identified by comparison with the human foamy virus (HFV). Unlike the HFV genomic organization, the SFV-1 pol gene overlaps the env gene; thus, the open reading frames reported for HFV between pol and env is not present in SFV-1. Comparisons of predicted amino acid sequences of HFV and SFV-1 reveal that the endonuclease domains of the pol genes are about 84% related. The region predicted to encode the SFV-1 extracellular env domain is 569 codons; SFV-1 and HFV have 64% amino acid similarity in this env domain. The predicted hydrophobic transmembrane env proteins of both HFV and SFV-1 show about 73% similarity. A total of 16 potential glycosylation sites are found in SFV-1 env, and 15 are found in HFV; 11 are shared. SFV-1 has 25 cysteine residues, and HFV has 23 residues; all 23 cysteine residues of HFV are conserved in SFV-1. This sequence analysis reveals that the human and simian foamy viruses are highly related.

Amino Acid Sequence↗

Bacteriophage T4 genome.

Phage T4 has provided countless contributions to the paradigms of genetics and biochemistry. Its complete genome sequence of 168,903 bp encodes about 300 gene products. T4 biology and its genomic sequence provide the best-understood model for modern functional genomics and proteomics. Variations on gene expression, including overlapping genes, internal translation initiation, spliced genes, translational bypassing, and RNA processing, alert us to the caveats of purely computational methods. The T4 transcriptional pattern reflects its dependence on the host RNA polymerase and the use of phage-encoded proteins that sequentially modify RNA polymerase; transcriptional activator proteins, a phage sigma factor, anti-sigma, and sigma decoy proteins also act to specify early, middle, and late promoter recognition. Posttranscriptional controls by T4 provide excellent systems for the study of RNA-dependent processes, particularly at the structural level. The redundancy of DNA replication and recombination systems of T4 reveals how phage and other genomes are stably replicated and repaired in different environments, providing insight into genome evolution and adaptations to new hosts and growth environments. Moreover, genomic sequence analysis has provided new insights into tail fiber variation, lysis, gene duplications, and membrane localization of proteins, while high-resolution structural determination of the "cell-puncturing device," combined with the three-dimensional image reconstruction of the baseplate, has revealed the mechanism of penetration during infection. Despite these advances, nearly 130 potential T4 genes remain uncharacterized. Current phage-sequencing initiatives are now revealing the similarities and differences among members of the T4 family, including those that infect bacteria other than Escherichia coli. T4 functional genomics will aid in the interpretation of these newly sequenced T4-related genomes and in broadening our understanding of the complex evolution and ecology of phages-the most abundant and among the most ancient biological entities on Earth.

Bacteriophage T4↗

The cobC gene of Salmonella typhimurium codes for a novel phosphatase involved in the assembly of the nucleotide loop of cobalamin.

We report the identification of a new locus, designated cobC, involved in the assembly of the nucleotide loop of cobalamin in Salmonella typhimurium. The cobC gene has been mapped, cloned, and sequenced. DNA sequence analysis suggested that cobC is divergently transcribed from the adjacent cobD gene and suggests that the regulatory region of these genes overlap. The cobC gene codes for a predicted polypeptide of 26 kDa with striking homology to phosphoglycerate mutase, fructose-2,6-bisphosphatase, and acid phosphatase enzymes. In vitro experiments demonstrated that CobC dephosphorylated the cobalamin biosynthetic intermediate N1-(5-phospho-alpha-D-ribosyl)-5,6-dimethylbenzimidazole to generate N1-alpha-D-ribosyl-5,6-dimethylbenzimidazole. In vivo data showed that the lack of cobC function blocks the synthesis of cobalamin from its precursors cobinamide and 5,6-dimethylbenzimidazole, i.e. it prevents the assembly of the nucleotide loop of cobalamin. Additionally, exogenous N1-alpha-D-ribosyl-5,6-dimethylbenzimidazole rescues the defect of a cobC mutant. We propose that cobC codes for a novel phosphatase whose primary role is in cobalamin biosynthesis. A model for the sequence of biosynthetic steps that assemble the nucleotide loop of cobalamin in S. typhimurium is presented.

Amino Acid Sequence↗

Generation and characterization of endonuclease G null mice.

Endonuclease G (endo G) is one of the most abundant nucleases in eukaryotic cells. It is encoded in the nucleus and imported to the mitochondrial intermembrane space. This nuclease is active on single- and double-stranded DNA. We genetically disrupted the endo G gene in mice without disturbing a conserved, overlapping gene of unknown function that is oriented tail to tail with the endo G gene. In these mice, the production of endo G protein is not detected, and the disruption abolishes the nuclease activity of endo G. The absence of endo G has no effect on mitochondrial DNA copy number, structure, or mutation rate over the first five generations. There is also no obvious effect on nuclear DNA degradation in standard apoptosis assays. The endo G null mice are viable and show no age-related or generational abnormalities anatomically or histologically. We infer that this highly conserved protein has no mitochondrial or apoptosis function that can discerned by the assays described here and that it may have a function yet to be determined. The early embryonic lethality of endo G null mice recently reported by others may be due to the disruption of the gene that overlaps the endo G gene.

Animals↗

Identification of chick rax/rx genes with overlapping patterns of expression during early eye and brain development.

We have isolated chick rax/rx cDNAs, cRaxL (chick Rax/Rx-like) and cRax, (chick Rax) and examined their expression patterns during early eye and brain development. The cRaxL cDNA encodes a 228 amino acid protein that is most closely related to the zebrafish Rx1 and Rx2. The cRax cDNA encodes a 317 amino acid protein, which shares higher homology with the Xenopus Rx. In addition to the homeodomain, the octapeptide and paired tail domains are conserved between the cRax and other vertebrate Rax/Rx, while cRaxL lacks the octapeptide containing N-terminal region which is conserved among all other members of the rax/rx gene family identified so far. The chick rax/rx genes are expressed in overlapping domains in the anterior neural ectoderm which corresponds to the forebrain and retina field, and later in the optic vesicle. cRax mRNA can be detected earlier than cRaxL prior to the formation of the notochord and its expression domain appears broader than that of cRaxL.

Animals↗

Performance evaluation of commercial short-oligonucleotide microarrays and the impact of noise in making cross-platform correlations.

BACKGROUND: Despite the widespread use of microarrays, much ambiguity regarding data analysis, interpretation and correlation of the different technologies exists. There is a considerable amount of interest in correlating results obtained between different microarray platforms. To date, only a few cross-platform evaluations have been published and unfortunately, no guidelines have been established on the best methods of making such correlations. To address this issue we conducted a thorough evaluation of two commercial microarray platforms to determine an appropriate methodology for making cross-platform correlations. RESULTS: In this study, expression measurements for 10,763 genes uniquely represented on Affymetrix U133A/B GeneChips and Amersham CodeLink UniSet Human 20 K microarrays were compared. For each microarray platform, five technical replicates, derived from the same total RNA samples, were labeled, hybridized, and quantified according to each manufacturers' standard protocols. The correlation coefficient (r) of differential expression ratios for the entire set of 10,763 overlapping genes was 0.62 between platforms. However, the correlation improved significantly (r = 0.79) when genes within noise were excluded. In addition to levels of inter-platform correlation, we evaluated precision, statistical-significance profiles, power, and noise levels for each microarray platform. Accuracy of differential expression was measured against real-time PCR for 25 genes and both platforms correlated well with r values of 0.92 and 0.79 for CodeLink and GeneChip, respectively. CONCLUSIONS: As a result of this study, we recommend using only genes called 'present' in cross-platform correlations. However, as in this study, a large number of genes may be lost from the correlation due to differing levels of noise between platforms. This is an important consideration given the apparent difference in sensitivity of the two platforms. Data from microarray analysis need to be interpreted cautiously and therefore, we provide guidelines for making cross-platform correlations. In all, this study represents the most comprehensive and specifically designed comparison of short-oligonucleotide microarray platforms to date using the largest set of overlapping genes.

Brain↗

Complete genome sequence of an Ebola virus (Sudan species) responsible for a 2000 outbreak of human disease in Uganda.

The entire genomic RNA of the Gulu (Uganda 2000) strain of Ebola virus was sequenced and compared to the genomes of other filoviruses. This data represents the first comprehensive genetic analysis for a representative isolate of the Sudan species of Ebola virus. The genome organization of the Sudan species is nearly identical to that of the Zaire species, but the presence of a gene overlap (between GP and VP30 genes) and a longer trailer sequence distinguish it from that of the Reston species. As has been observed with other filoviruses, stemloop structures were predicted to form at the 5' end of Ebola Sudan mRNA molecules, and the genomic RNA termini showed a high degree of sequence complimentarity. Comparisons of the amino acid sequences of encoded gene products shows that there is a comparable level of identity or similarity between Ebola virus species, with Sudan and Zaire actually showing a slightly closer relationship to the Reston species than to one another. These comparisons also indicated that the VP24 is the most conserved Ebola virus protein (followed closely by the VP40 and L proteins), while the GP is the least conserved gene product. The most divergent regions were seen in the C-terminus of GP1 (mucin-like region) and within the C-terminal third of the nucleoprotein sequence.

5' Untranslated Regions↗

Evolutionary changes of nucleotide sequences of papova viruses BKV and SV40: they are possibly hybrids.

Complete nucleotide sequences were compared between papova viruses BKV and SV40 and the degrees of sequence divergences were compared between structurally and/or functionally different segments or genes in details. It was shown that the rate of synonymous substitution is not only very high but also approximately uniform among different genes in these viruses as in eukaryotic genes examined to date. While all the non-coding regions including the intron showed marked sequence preservation which is in sharp contrasted with the case of eukaryotic genes where the large bulk of non-coding regions evolve at a rate as rapidly as that of synonymous substitution. It is remarkable that a long continuous stretch of sequence including the putative VPX gene and a 5' half of VP2 gene showed strong homology between BKV and SV40. A close examination of the pattern of base substitutions revealed that this unusual homology was derived by recombination between the two viruses during their evolution. On the basis of the pattern of base substitutions and the bias in code word utilization, we also showed that the putative VPX gene actually could code for a functional polypeptide. In papova viruses, the 3' terminal sequence of VP2/3 gene overlaps with the 5' terminal sequence of VPI gene. The pattern of base substitutions in the overlapping segment was examined in detail in comparison with those in the non-overlapping portions of VP2/3 and VP1 genes. It was shown that the evolutionary mode of the overlapping genes is in good agreement with our previous prediction.

BK Virus↗

Resolving the functions of overlapping viral genes by site-specific mutagenesis at a mRNA splice site.

Early region IA of human adenoviruses encodes a function required for normal induction of early viral genes and virus-induced cell transformation. The region is expressed at early times as two overlapping spliced mRNAs, 12S and 13S, which encode closely related proteins. To distinguish between the functions of these proteins, a single T leads to G transversion was constructed which prevents splicing of the 12S mRNA. This transversion, in the second base of the 12S mRNA intron, does not alter the protein encoded by the 13S mRNA due to degeneracy in the genetic code. Studies with this mutant demonstrated that only the 13S mRNA encodes the regulatory protein required for normal early gene expression.

Adenoviruses, Human↗