Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Natively disordered proteins: functions and predictions.

Proteins can exist in at least three forms: the ordered form (solid-like), the partially folded form (collapsed, molten globule-like or liquid-like) and the extended form (extended, random coil-like or gas-like). The protein trinity hypothesis has two components: (i) a given native protein can be in any one of the three forms, depending on the sequence and the environment; and (ii) function can arise from any one of the three forms or from transitions between them. In this study, bioinformatics and data mining were used to investigate intrinsic disorder in proteins and develop neural network-based predictors of natural disordered regions (PONDR) that can discriminate between ordered and disordered residues with up to 84% accuracy. Predictions of intrinsic disorder indicate that the three kingdoms follow the disorder ranking eubacteria < archaebacteria << eukaryotes, with approximately half of eukaryotic proteins predicted to contain substantial regions of intrinsic disorder. Many of the known disordered regions are involved in signalling, regulation or control. Involvement of highly flexible or disordered regions in signalling is logical: a flexible sensor more readily undergoes conformational change in response to environmental perturbations than does a rigid one. Thus, the increased disorder in the eukaryotes is likely the direct result of an increased need for signalling and regulation in nucleated organisms. PONDR can also be used to detect molecular recognition elements that are disordered in the unbound state and become structured when bound to a biologically meaningful partner. Application of disorder predictions to cell-signalling, cancer-associated and control protein databases supports the widespread occurrence of protein disorder in these processes.

Amino Acid Sequence↗

Expression profiling and bioinformatic analyses of a novel stress-regulated multispanning transmembrane protein family from cereals and Arabidopsis.

Cold acclimation is a multigenic trait that allows hardy plants to develop efficient tolerance mechanisms needed for winter survival. To determine the genetic nature of these mechanisms, several cold-responsive genes of unknown function were identified from cold-acclimated wheat (Triticum aestivum). To identify the putative functions and structural features of these new genes, integrated genomic approaches of data mining, expression profiling, and bioinformatic predictions were used. The analyses revealed that one of these genes is a member of a small family that encodes two distinct groups of multispanning transmembrane proteins. The cold-regulated (COR)413-plasma membrane and COR413-thylakoid membrane groups are potentially targeted to the plasma membrane and thylakoid membrane, respectively. Further sequence analysis of the two groups from different plant species revealed the presence of a highly conserved phosphorylation site and a glycosylphosphatidylinositol-anchoring site at the C-terminal end. No homologous sequences were found in other organisms suggesting that this family is specific to the plant kingdom. Intraspecies and interspecies comparative gene expression profiling shows that the expression of this gene family is correlated with the development of freezing tolerance in cereals and Arabidopsis. In addition, several members of the family are regulated by water stress, light, and abscisic acid. Structure predictions and comparative genome analyses allow us to propose that the cor413 genes encode putative G-protein-coupled receptors.

Acclimatization↗

Statistical and visual morph movie analysis of crystallographic mutant selection bias in protein mutation resource data.

Structural studies of the effects of non-silent mutations on protein conformational change are an important key in deciphering the language that relates protein amino acid primary structure to tertiary structure. Elsewhere, we presented the Protein Mutant Resource (PMR) database, a set of online tools that systematically identified groups of related mutant structures in the Protein DataBank (PDB), accurately inferred mutant classifications in the Gene Ontology using an innovative, statistically rigorous data-mining algorithm with more general applicability, and illustrated the relationship of these mutant structures via an intuitive user interface. Here, we perform a comprehensive statistical analysis of the effect of PMR mutations on protein tertiary structure. We find that, although the PMR does contain spectacular examples of conformational change, in general there is a counter-intuitive inverse relationship between conformational change (measured as C-alpha displacement or RMS of the core structure) and the number of mutations in a structure. That is, point mutations by structural biologists present in the PDB contrast naturally evolved mutations. We compare the frequency of mutations in the PMR/PDB datasets against the accepted PAM250 natural amino acid mutation frequency to confirm these observations. We generated morph movies from PMR structure pairs using technology previously developed for the Macromolecular Motions Database (http://molmovdb.org), allowing bioinformaticians, geneticists, protein engineers, and rational drug designers to analyze visually the mechanisms of protein conformational change and distinguish between conformational change due to motions (e.g., ligand binding) and mutations. The PMR morph movies and statistics can be freely viewed from the PMR website, http://pmr.sdsc.edu.

Amino Acid Sequence↗

A microsatellite-based, gene-rich linkage map for the AA genome of Arachis (Fabaceae).

Cultivated peanut (Arachis hypogaea) is an important crop, widely grown in tropical and subtropical regions of the world. It is highly susceptible to several biotic and abiotic stresses to which wild species are resistant. As a first step towards the introgression of these resistance genes into cultivated peanut, a linkage map based on microsatellite markers was constructed, using an F(2) population obtained from a cross between two diploid wild species with AA genome (A. duranensis and A. stenosperma). A total of 271 new microsatellite markers were developed in the present study from SSR-enriched genomic libraries, expressed sequence tags (ESTs), and by "data-mining" sequences available in GenBank. Of these, 66 were polymorphic for cultivated peanut. The 271 new markers plus another 162 published for peanut were screened against both progenitors and 204 of these (47.1%) were polymorphic, with 170 codominant and 34 dominant markers. The 80 codominant markers segregating 1:2:1 (P<0.05) were initially used to establish the linkage groups. Distorted and dominant markers were subsequently included in the map. The resulting linkage map consists of 11 linkage groups covering 1,230.89 cM of total map distance, with an average distance of 7.24 cM between markers. This is the first microsatellite-based map published for Arachis, and the first map based on sequences that are all currently publicly available. Because most markers used were derived from ESTs and genomic libraries made using methylation-sensitive restriction enzymes, about one-third of the mapped markers are genic. Linkage group ordering is being validated in other mapping populations, with the aim of constructing a transferable reference map for Arachis.

Arachis↗

Cold-regulated cereal chloroplast late embryogenesis abundant-like proteins. Molecular characterization and functional analyses.

Cold acclimation and freezing tolerance are the result of complex interaction between low temperature, light, and photosystem II (PSII) excitation pressure. Previous results have shown that expression of the Wcs19 gene is correlated with PSII excitation pressure measured in vivo as the relative reduction state of PSII. Using cDNA library screening and data mining, we have identified three different groups of proteins, late embryogenesis abundant (LEA) 3-L1, LEA3-L2, and LEA3-L3, sharing identities with WCS19. These groups represent a new class of proteins in cereals related to group 3 LEA proteins. They share important characteristics such as a sorting signal that is predicted to target them to either the chloroplast or mitochondria and a C-terminal sequence that may be involved in oligomerization. The results of subcellular fractionation, immunolocalization by electron microscopy and the analyses of target sequences within the Wcs19 gene are consistent with the localization of WCS19 within the chloroplast stroma of wheat (Triticum aestivum) and rye (Secale cereale). Western analysis showed that the accumulation of chloroplastic LEA3-L2 proteins is correlated with the capacity of different wheat and rye cultivars to develop freezing tolerance. Arabidopsis was transformed with the Wcs19 gene and the transgenic plants showed a significant increase in their freezing tolerance. This increase was only evident in cold-acclimated plants. The putative function of this protein in the enhancement of freezing tolerance is discussed.

Acclimatization↗

Insertion-deletion polymorphisms in 3' regions of maize genes occur frequently and can be used as highly informative genetic markers.

Single-nucleotide polymorphisms (SNPs) are the most frequent variations in the genome of any organism. SNP discovery approaches such as resequencing or data mining enable the identification of insertion deletion (indel) polymorphisms. These indels can be treated as biallelic markers and can be utilized for genetic mapping and diagnostics. In this study 655 indels have been identified by resequencing 502 maize (Zea mays) loci across 8 maize inbreds (selected for their high allelic variation). Of these 502 loci, 433 were polymorphic, with indels identified in 215 loci. Of the 655 indels identified, single-nucleotide indels accounted for more than half (54.8%) followed by two- and three-nucleotide indels. A high frequency of 6-base (3.4%) and 8-base (2.3%) indels were also observed. When analysis is restricted to the B73 and Mol7 genotypes, 53% of the loci analyzed contained indels, with 42% having an amplicon size difference. Three novel miniature inverted-repeat transposable element (MITE)-like sequences were identified as insertions near genes. The utility of indels as genetic markers was demonstrated by using indel polymorphisms to map 22 loci in a B73 x Mo17 recombinant inbred population. This paper clearly demonstrates that the resequencing of 3' EST sequence and the discovery and mapping of indel markers will position corresponding expressed genes on the genetic map.

Chromosome Mapping↗

Diversity and evolution of Bdellovibrio-and-like organisms (BALOs), reclassification of Bacteriovorax starrii as Peredibacter starrii gen. nov., comb. nov., and description of the Bacteriovorax-Peredibacter clade as Bacteriovoracaceae fam. nov.

A phylogenetic analysis of Bdellovibrio-and-like organisms (BALOs) was performed. It was based on the characterization of 71 strains and on all consequent 16S rRNA gene sequences available in databases, including clones identified by data-mining, totalling 120 strains from very varied biotopes. Amplified rDNA restriction analysis (ARDRA) accurately reflected the diversity and phylogenetic affiliation of BALOs, thereby providing an efficient screening tool. Extensive phylogenetic analysis of the 16S rRNA gene sequences revealed great diversity within the Bdellovibrio (> 14 % divergence) and Bacteriovorax (> 16 %) clades, which comprised nine and eight clusters, respectively, exhibiting more than 3 % intra-cluster divergence. The clades diverged by more than 20 %. The analysis of conserved 16S rRNA secondary structures showed that Bdellovibrio contained motifs atypical of the delta-Proteobacteria, suggesting that it is ancestral to Bacteriovorax. While none of the Bdellovibrio strains were of marine origin, Bacteriovorax included separate soil/freshwater and marine-specific groups. On the basis of their extensive diversity and the large distance separating the groups, it is proposed that Bacteriovorax starrii be placed into a new genus, Peredibacter gen. nov., with Peredibacter starrii A3.12T (= ATCC 15145T = NCCB 72004T) as its type strain. Also proposed is a redefinition of the Bdellovibrio and the Bacteriovorax-Peredibacter lineages as two different families, i.e. 'Bdellovibrionaceae' and a new family, Bacteriovoracaceae. Also, a re-evaluation of oligonucleotides targeting BALOs is presented, and the implications of the large diversity of these organisms and of their distribution in very different environments are discussed.

Bdellovibrio↗

Characterization of a rice class II metallothionein gene: tissue expression patterns and induction in response to abiotic factors.

Data mining the complete rice genome sequences revealed a genomic fragment encoding a characteristic metallothionein (MT) protein, and its full-length cDNA was isolated from rice developing seeds by RT-PCR. This cDNA, designated OsMT-II-1a, contains an open reading frame of 264 bp encoding a protein of 87 amino acid residues. The predicted amino acid sequence was shown to have structural features characteristic of plant class II MT proteins. By sequence analysis of its 5'-flanking region, one putative TATA box, four putative CAAT boxes, and several short sequences homologous to previously reported regulatory cis-elements were identified. Northern blot analysis showed that accumulation of OsMT-II-1a mRNA is specifically abundant in developing seeds and 2-day glumes after pollination, and OsMT-II-1a transcription can markedly be induced by H2O2, paraquat, SNP, ethephon, ABA and SA, but barely by metal ions or other exogenous abiotic factors such as low temperature and PEG. These results coincide with the prediction of existing regulatory cis-elements in its 5'-flanking region. Taken together, the above results suggest that the processes of pollination and seed development might be mediated, at least in part, by expression of the OsMT-II-1a gene that is regulated by several abiotic factors.

Amino Acid Sequence↗

Cloning of a novel neuronally expressed orphan G-protein-coupled receptor which is up-regulated by erythropoietin, interacts with microtubule-associated protein 1b and colocalizes with the 5-hydroxytryptamine 2a receptor.

G-protein-coupled receptors (GPCRs) are the largest group of cell surface molecules involved in signal transduction and are receptors for a wide variety of stimuli ranging from light, calcium and odourants to biogenic amines and peptides. It is assumed that systematic genomic data-mining has identified the overwhelming majority of all remaining GPCRs in the genome. Here we report the cloning of a novel orphan GPCR which was identified in a search for erythropoietin-induced genes in the brain as a strongly up-regulated gene. This unknown gene coded for a protein which had a seven-transmembrane topology and key features typical of GPCRs of the A family but a low overall identity to all known GPCRs. The protein, coded ee3, has an unusually high evolutionary conservation and is expressed in neurons in diverse areas of the CNS with relation to integrative functions or motor tasks. A yeast two-hybrid screen for interacting proteins revealed binding to the microtubule-associated protein (MAP) 1b. Coupling to MAP1a has been described for another cognate GPCR, the 5-hydroxytryptamine (5HT) 2a receptor. Surprisingly, we found complete colocalization of ee3 and the 5HT2a receptor. The interaction with MAP1b proved to be critical for the stability or folding of ee3 as in mice lacking MAP1b the ee3 protein was undetectable by immunohistochemistry, although messenger RNA levels remained unchanged. We propose that ee3 is a highly interesting new orphan GPCR with potential connections to erythropoietin and 5HT2a receptor signalling.

Amino Acid Sequence↗

Genomic characterization of Rim2/Hipa elements reveals a CACTA-like transposon superfamily with unique features in the rice genome.

The availability of huge amounts of rice genome sequence now permits large-scale analysis of the structure and molecular characteristics of the previously identified transposase-encoding Rim2 (also called Hipa) element, which is transcriptionally activated by infection with the fungal pathogen Magnaporthe grisea and by treatment with the corresponding fungal elicitor. Based on genomic cloning and data mining from 230 Mb of rice genome sequence, 347 Rim2 elements, with an average size of 5.8 kb, were identified. This indicates that an estimated total of 600-700 Rim2 elements are present in the whole genome. Rim2 insertions occur non-randomly on the chromosomes, as visualized by fluorescence in situ hybridization. The elements harbor 16-bp terminal inverted repeats with the core sequence CACTG, 16-bp sub-terminal repeats, internal variable regions, 3-bp target sequence duplications in the flanking regions, and genes coding for Rim2 proteins (the putative transposase) and hydroxyproline-rich glycoproteins. High levels of insertion into genic regions are observed for members of this family, and the transposition history of the family can be deduced from the high level of shared sequences and analysis of repeat target sites of the elements. Phylogenetic analysis indicates that the putative RIM2 proteins fall into a subgroup distinct from the TNP2-like subgroup of transposases. Southern hybridization with genomic DNA from monocotyledonous and dicotyledonous plants demonstrates that the RIM2-coding sequence is unique to the Oryza genome. Our results demonstrate that the Rim2 elements from rice belong to a distinct superfamily of CACTA-like elements with evolutionary diversity.

Amino Acid Sequence↗

Identification of OCT6 as a novel organic cation transporter preferentially expressed in hematopoietic cells and leukemias.

OBJECTIVE: Human organic cation transporters (OCTs) play a critical role in the cellular uptake and efflux of endogenous cationic substrates and hydrophilic exogenous xenobiotics. We sought to identify OCT genes preferentially expressed in hematopoietic cells. MATERIALS AND METHODS: We isolated a novel OCT, named OCT6, by data-mining human expressed sequence tag databases for sequences homologous to known OCT genes. We developed a quantitative reverse transcriptase polymerase chain reaction assay to determine the relative expression of this gene in 50 cancer cell lines and in tissues. RESULTS: The two highest expressing cell lines were the leukemia cell lines HL-60 and MOLT4. Quantitative reverse transcriptase polymerase chain reaction analysis using a normal tissue cDNA panel demonstrated that this transport gene is highly expressed in testis and fetal liver, with detectable RNA levels in bone marrow and peripheral blood leukocytes. Unlike other OCT genes, RNA levels were not detectable in placenta, liver, or kidney. To further define the expression of OCT6 in hematopoietic tissues, we measured OCT6 RNA levels in sorted peripheral blood cell populations and found a clear enrichment of OCT6-expressing cells in purified CD34(+) cells. To determine if OCT6 was highly expressed in leukemias, we examined circulating leukemia cells from 25 patients and found high levels of OCT6 RNA in all specimens in comparison with liver, kidney, and placenta. CONCLUSIONS: The results demonstrate the existence of a novel OCT preferentially expressed in human hematopoietic tissues, including CD34(+) cells and leukemia cells. Its narrow tissue distribution, potential for substrate specificity, and close homology to other cell membrane transporters make OCT6 an attractive target for the treatment of leukemia.

Amino Acid Sequence↗

A genome-wide identification of E2F-regulated genes in Arabidopsis.

The completion of the Arabidopsis genomic sequence offers the possibility to extract global information about regulatory mechanisms. Here, we describe a data mining strategy in combination with gene expression analysis to identify bona fide genes regulated by the E2F transcription factor. Starting with a genome-wide search of chromosomal sites containing E2F-binding sites, we studied in depth two of the most abundant E2F-binding sites within the Arabidopsis genome and identified over 180 potential E2F target genes. Among them and in addition to cell cycle-related genes, we have also identified genes belonging to other functional categories, e.g. transcription, stress and defense or signaling. We have determined the expression levels of genes selected from different categories under two experimental situations. Using cultured cells partially synchronized with aphidicolin, we found that most potential E2F targets identified in silico show a cell cycle-regulated expression pattern with a peak in early/mid S-phase. In addition, we used Arabidopsis transgenic plants expressing a DP gene containing a truncated DNA-binding domain, which likely has a dominant-negative effect on AtE2Fa, b and c (also named AtE2F3, 1 and 2, respectively), which require DP for efficient DNA binding. Contrary to the up-regulation observed in early/mid S-phase-cultured cells, the expression of a large number of potential E2F targets was decreased in the transgenic plants. Our results strongly support that the RBR/E2F pathway plays a crucial role in regulating the expression of the genes identified in this study.

Amino Acid Sequence↗

Bioinformatics-based discovery of a novel factor with apparent specificity to colon cancer.

In a previous study, a data mining tool called Digital Differential Display (DDD) from the Cancer Genome Anatomy Project (CGAP) was used to predict solid tumor- and organ-specific genes from the expressed sequence tag (EST) database. To validate the use of bioinformatics approaches in gene discovery, one of the ESTs, which was predicted to be colon tumor-specific, was chosen for further study. Reverse Transcriptase-Polymerase Chain Reaction (RT-PCR) analysis of matched sets of cDNAs from normal and colon tumor tissues indicated that the EST was specifically expressed in the majority of colon tumors. Expression was also detected in early adenomas. Among other normal tissues, EST expression was detected only in the small intestine. The colon tumor specificity of this EST was inferred from the lack of expression in carcinomas of the breast, lung, ovary, pancreas and prostate. To validate the computational prediction of specificity, a full-length cDNA encompassing the entire open reading frame was cloned and, in view of its apparent specificity to the colon tumors, this gene was termed Colon Carcinoma Related Gene (CCRG). CCRG encodes a novel cysteine-rich motif and a putative signal peptide sequence. Supernatant from COS cells transfected with the CCRG expression vector stimulated proliferation of colon cancer cells. Immunoreactive CCRG was also detected in the paraffin sections of colon tumor samples. CCRG belongs to a new class of growth factors and may be important in the diagnosis and treatment of colon cancers. Identification of CCRG using bioinformatics approaches validates gene discovery using computational approaches.

Amino Acid Sequence↗

The gene ncgl2918 encodes a novel maleylpyruvate isomerase that needs mycothiol as cofactor and links mycothiol biosynthesis and gentisate assimilation in Corynebacterium glutamicum.

Data mining of the Corynebacterium glutamicum genome identified 4 genes analogous to the mshA, mshB, mshC, and mshD genes that are involved in biosynthesis of mycothiol in Mycobacterium tuberculosis and Mycobacterium smegmatis. Individual deletion of these genes was carried out in this study. Mutants mshC- and mshD- lost the ability to produce mycothiol, but mutant mshB- produced mycothiol as the wild type did. The phenotypes of mutants mshC- and mshD- were the same as the wild type when grown in LB or BHIS media, but mutants mshC- and mshD- were not able to grow in mineral medium with gentisate or 3-hydroxybenzoate as carbon sources. C. glutamicum assimilated gentisate and 3-hydroxybenzoate via a glutathione-independent gentisate pathway. In this study it was found that the maleylpyruvate isomerase, which catalyzes the conversion of maleylpyruvate into fumarylpyruvate in the glutathione-independent gentisate pathway, needed mycothiol as a cofactor. This mycothiol-dependent maleylpyruvate isomerase gene (ncgl2918) was cloned, actively expressed, and purified from Escherichia coli. The purified mycothiol-dependent isomerase is a monomer of 34 kDa. The apparent Km and Vmax values for maleylpyruvate were determined to be 148.4 +/- 11.9 microM and 1520 +/- 57.4 micromol/min/mg, respectively (mycothiol concentration, 2.5 microM). Previous studies had shown that mycothiol played roles in detoxification of oxidative chemicals and antibiotics in streptomycetes and mycobacteria. To our knowledge, this is the first demonstration that mycothiol is essential for growth of C. glutamicum with gentisate or 3-hydroxybenzoate as carbon sources and the first characterization of a mycothiol-dependent maleylpyruvate isomerase.

Amino Acid Sequence↗

Statistical characterization of the charge state and residue dependence of low-energy CID peptide dissociation patterns.

Data mining was performed on 28 330 unique peptide tandem mass spectra for which sequences were assigned with high confidence. By dividing the spectra into different sets based on structural features and charge states of the corresponding peptides, chemical interactions involved in promoting specific cleavage patterns in gas-phase peptides were characterized. Pairwise fragmentation maps describing cleavages at all Xxx-Zzz residue combinations for b and y ions reveal that the difference in basicity between Arg and Lys results in different dissociation patterns for singly charged Arg- and Lys-ending tryptic peptides. While one dominant protonation form (proton localized) exists for Arg-ending peptides, a heterogeneous population of different protonated forms or more facile interconversion of protonated forms (proton partially mobile) exists for Lys-ending peptides. Cleavage C-terminal to acidic residues dominates spectra from singly charged peptides that have a localized proton and cleavage N-terminal to Pro dominates those that have a mobile or partially mobile proton. When Pro is absent from peptides that have a mobile or partially mobile proton, cleavage at each peptide bond becomes much more prominent. Whether the above patterns can be found in b ions, y ions, or both depends on the location of the proton holder(s) in multiply protonated peptides. Enhanced cleavages C-terminal to branched aliphatic residues (Ile, Val, Leu) are observed in both b and y ions from peptides that have a mobile proton, as well as in y ions from peptides that have a partially mobile proton; enhanced cleavages N-terminal to these residues are observed in b ions from peptides that have a partially mobile proton. Statistical tools have been designed to visualize the fragmentation maps and measure the similarity between them. The pairwise cleavage patterns observed expand our knowledge of peptide gas-phase fragmentation behaviors and may be useful in algorithm development that employs improved models to predict fragment ion intensities.

Amino Acid Sequence↗

Comparative sequence and expression analyses of four mammalian VPS4 genes.

The VPS4 gene is a member of the AAA-family; it codes for an ATPase which is involved in lysosomal/endosomal membrane trafficking. VPS4 genes are present in virtually all eukaryotes. Exhaustive data mining of all available genomic databases from completely or partially sequenced organisms revealed the existence of up to three paralogues, VPS4a, -b, and -c. Whereas in the genome of lower eukaryotes like yeast only one VPS4 representative is present, we found that mammals harbour two paralogues, VPS4a and VPS4b. Most interestingly, the Fugu fish contains a third VPS4 paralogue (VPS4c). Sequence comparison of the three VPS4 paralogues indicates that the Fugu VPS4c displays sequence features intermediate between VPS4a and VPS4b. Using complete mammalian VPS4a and VPS4b cDNA clones as probes, genomic clones of both VPS4 paralogues in human and mouse were identified and sequenced. The chromosomal loci of all four VPS4 genes were determined by independent methods. A BLAST search of the human genome database with the human VPS4A sequence yielded a double match, most likely due to a faulty assembly of sequence contigs in the human draft sequence. Fluorescent in situ hybridization and radiation hybrid analyses demonstrated that human and mouse VPS4A/a and VPS4B/b are located on syntenic chromosomal regions. Northern blot and semi-quantitative reverse transcription analyses showed that mouse VPS4a and VPS4b are differentially expressed in different organs, suggesting that the two paralogues have developed different functional properties since their divergence. To investigate the subcellular distribution of the murine VPS4 paralogues, we transiently expressed various fluorescent VPS4 fusion proteins in mouse 3T3 cells. All tested VPS4 fusion proteins were found in the cytosol. Expression of dominant-negative mutant VPS4 fusion proteins led to their concentration in the perinuclear region. Co-expression of VPS4a-GFP and VPS4b-dsRed fusion proteins revealed a partial co-localization that was most prominent with mutant VPS4a and VPS4b proteins. A physical interaction between the mouse paralogues was also supported by two-hybrid analyses.

3T3 Cells↗

Identification of a new pebp2alphaA2 isoform from zebrafish runx2 capable of inducing osteocalcin gene expression in vitro.

UNLABELLED: The zebrafish runx2b transcription factor is an ortholog of RUNX2 and is highly conserved at the structural level. The runx2b pebp2alphaA2 isoform induces osteocalcin gene expression by binding to a specific region of the promoter and seems to have been selectively conserved in the teleost lineage. INTRODUCTION: RUNX2 (also known as CBFA1/Osf2/AML3/PEBP2alphaA) is a transcription factor essential for bone formation in mammals, as well as for osteoblast and chondrocyte differentiation, through regulation of expression of several bone- and cartilage-related genes. Since its discovery, Runx2 has been the subject of intense studies, mainly focused in unveiling regulatory targets of this transcription factor in high vertebrates. However, no single study has been published addressing the role of Runx2 in bone metabolism of low vertebrates. While analyzing the zebrafish (Danio rerio) runx2 gene, we identified the presence of two orthologs of RUNX2, which we named runx2a and runx2b and cloned a pebp2alphaA-like transcript of the runx2b gene, which we named pebp2alphaA2. MATERIALS AND METHODS: Zebrafish runx2b gene and cDNA were isolated by RT-PCR and sequence data mining. The 3D structure of runx2b runt domain was modeled using mouse Runx1 runt as template. The regulatory effect of pebp2alphaA2 on osteocalcin expression was analyzed by transient co-transfection experiments using a luciferase reporter gene. Phylogenetic analysis of available Runx sequences was performed with TREE_PUZZLE 5.2. and MrBayes. RESULTS AND CONCLUSIONS: We showed that the runx2b gene structure is highly conserved between mammals and fish. Zebrafish runx2b has two promoter regions separated by a large intron. Sequence analysis suggested that the runx2b gene encodes three distinct isoforms, by a combination of alternative splicing and differential promoter activation, as described for the human gene. We have cloned a pebp2alphaA-like transcript of the runx2b gene, which we named pebp2alphaA2, and showed its high degree of sequence similarity with the mammalian pebp2alphaA. The cloned zebrafish osteocalcin promoter was found to contain three putative runx2-binding elements, and one of them, located at -221 from the ATG, was capable of mediating pebp2alphaA2 transactivation. In addition, cross-species transactivation was also confirmed because the mouse Cbfa1 was able to induce the zebrafish osteocalcin promoter, whereas the zebrafish pebp2alphaA2 activated the murine osteocalcin promoter. These results are consistent with the high degree of evolutionary conservation of these proteins. The 3D structure of the runx2b runt domain was modeled based on the runt domain of mouse Runx1. Results show a high degree of similarity in the 3D configuration of the DNA binding regions from both domains, with significant differences only observed in non-DNA binding regions or in DNA-binding regions known to accommodate considerable structure flexibility. Phylogenetic analysis was used to clarify the relationship between the isoforms of each of the two zebrafish Runx2 orthologs and other Runx proteins. Both zebrafish runx2 genes clustered with other Runx2 sequences. The duplication event seemed, however, to be so old that, whereas Runx2b clearly clusters with the other fish sequences, it is unclear whether Runx2a clusters with Runx2 from higher vertebrates or from other fish.

Amino Acid Sequence↗

Molecular cloning and primary functional analysis of a novel human testis-specific gene.

In this study, a new data mining tool called Digital Differential Display (DDD) from the NCBI was used to predict testis-specific expressed genes from the expressed sequence tag (EST) database. DDD (digital differential display) was performed between nine testis libraries and seventy libraries derived from other tissues. We identified a new contig of ESTs (HS. 326528) which was from testis libraries. To validate the use of bioinformatic approaches in gene discovery, the ESTs (HS. 326528), which were predicted to be testis-specific, were chosen for further study. Reverse Transcriptase-Polymerase Chain Reaction (RT-PCR) analysis of matched sets of cDNAs from testis and other tissues indicated that the ESTs were specifically expressed in testis. This result was further validated by multi-tissue Northern blot. The full-length cDNA encompassing the entire open reading frame was cloned and, in view of its apparent specificity to testis, the gene was termed homo sapiens spermatogenesis-related gene 8---SRG8 (GenBank accession number: AY489187). The gene whose full cDNA length is 1 044 bp containing 3 exons and 2 introns is located in human chromosome 15q26.2, the cDNA encodes a novel protein of 105 amino acides with a theoretical molecular weight of 11.7 kD and isoelectric point of 10.09 which shares no significant homology with any known proteins in database. Real time PCR analysis of testis of different developmental periods revealed that SRG8 gene is significantly expressed in adult testis. The green fluorescent protein produced by pEGFP-C3/SRG8 was detected in the nucleus of HeLa cells after 24 h post-transfection. Cell cycle analysis showed that SRG8 can accelerate HeLa cells to traverse the S-phase and enter the G2-phase compared with the control without transfection of SRG8, which suggested that this gene plays an important role in the development of testis. The discovery of SRG8 showed that DDD combined with experiments is a feasible, time-saving strategy to identify new candidate genes for testis-specific development

Adult↗