Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Usage of putative chicken U6 promoters for vector-based RNA interference.

Gene silencing with short interfering RNA (siRNA) expression vectors is a powerful method for the analysis of gene functions. For the expression of siRNA in mammalian cells, mammalian U6 small nuclear RNA (snRNA) promoters are widely used. However, the mammalian U6 promoter might not function well in other species. In this study, we cloned four putative chicken U6 promoters by PCR and analyzed their functions. First, we screened the chicken genomic database using the human U6 snRNA gene and identified four candidate sequences. The sequences contained some control elements in their promoter regions, but as we could not rule out that they were pseudogenes, we amplified these sequences and used them as promoters for short hairpin RNA (shRNA) expression. Using the firefly luciferase (Luc) gene as a target, transient expression assays were performed with chicken ovary-derived cells. All four putative chicken U6 promoters exhibited suppressive activity toward Luc, and so could act as a promoter for expression of the snRNA gene in the chicken genome. The promoter activity was not as strong as that of a commercially available siRNA expression vector. This probably reflects artificial sequences between the promoters and synthetic DNA encoding shRNA.

Animals↗

Molecular phylogeny of the kelch-repeat superfamily reveals an expansion of BTB/kelch proteins in animals.

BACKGROUND: The kelch motif is an ancient and evolutionarily-widespread sequence motif of 44-56 amino acids in length. It occurs as five to seven repeats that form a beta-propeller tertiary structure. Over 28 kelch-repeat proteins have been sequenced and functionally characterised from diverse organisms spanning from viruses, plants and fungi to mammals and it is evident from expressed sequence tag, domain and genome databases that many additional hypothetical proteins contain kelch-repeats. In general, kelch-repeat beta-propellers are involved in protein-protein interactions, however the modest sequence identity between kelch motifs, the diversity of domain architectures, and the partial information on this protein family in any single species, all present difficulties to developing a coherent view of the kelch-repeat domain and the kelch-repeat protein superfamily. To understand the complexity of this superfamily of proteins, we have analysed by bioinformatics the complement of kelch-repeat proteins encoded in the human genome and have made comparisons to the kelch-repeat proteins encoded in other sequenced genomes. RESULTS: We identified 71 kelch-repeat proteins encoded in the human genome, whereas 5 or 8 members were identified in yeasts and around 18 in C. elegans, D. melanogaster and A. gambiae. Multiple domain architectures were identified in each organism, including previously unrecognised forms. The vast majority of kelch-repeat domains are predicted to form six-bladed beta-propellers. The most prevalent domain architecture in the metazoan animal genomes studied was the BTB/kelch domain organisation and we uncovered 3 subgroups of human BTB/kelch proteins. Sequence analysis of the kelch-repeat domains of the most robustly-related subgroups identified differences in beta-propeller organisation that could provide direction for experimental study of protein-binding characteristics. CONCLUSION: The kelch-repeat superfamily constitutes a distinct and evolutionarily-widespread family of beta-propeller domain-containing proteins. Expansion of the family during the evolution of multicellular animals is mainly accounted for by a major expansion of the BTB/kelch domain architecture. BTB/kelch proteins constitute 72 % of the kelch-repeat superfamily of H. sapiens and form three subgroups, one of which appears the most-conserved during evolution. Distinctions in propeller blade organisation between subgroups 1 and 2 were identified that could provide new direction for biochemical and functional studies of novel kelch-repeat proteins.

Amino Acid Motifs↗

Identification of sigmaB-dependent promoters using consensus-directed search of Streptomyces coelicolor genome.

SigmaB plays an important role in both osmoprotection and proper differentiation in Streptomyces coelicolor A3(2). We searched for candidate members of the sigmaB regulon from the genome database, using the consensus promoter sequence (GNNTN14-16GGGTAC/T). The list consists of 115 genes, and includes all the known sigmaB target genes and many other genes whose functions are related to stress protection and differentiation.

Adaptation, Physiological↗

Defining relationships between the known members of the cytochrome P450 3A subfamily, including five putative chimpanzee members.

An analysis of the cytochrome P450 3A subfamily (CYP3A) was undertaken in order to define relationships across species among subfamily members. Some members were excluded due to incomplete sequences, while others were held in abeyance because of their almost complete homology. This is the first publication of five chimpanzee CYP3A genes-CYP3A4, CYP3A5, CYP3A7, CYP3A43, and CYP3A67. This project utilized two approaches for characterizing possible relationships-phylogenetic analysis and genomic structure. For the phylogenetic analysis, both nucleotide and amino acid sequences were aligned in silico using the CLUSTAL algorithm, and then visually inspected for accuracy. Three different computer software packages were utilized: MEGA 2.1, TREECON 1.3b, and PHYLIP 3.5. Multiple methods were used: neighbor-joining (NJ), minimum evolution (ME), maximum parsimony (MP), and maximum likelihood (ML). The resulting topologies were compared against each other to define the consensus topology. In addition, the chimpanzee, human, mouse, and rat genome databases were searched for intron/exon information pertaining to the included genes. Both methods suggest the same conclusion, defining orthologs is plausible between similar species (i.e., mouse and rat), but is less useful between species of different orders (i.e., primate and rodent) or classes (i.e., mammal and avian).

Animals↗

[Molecular cloning and sequence analyzing of cytoplasmic ribosomal protein gene OsRPS7 from rice (Oryza sativa)].

Using the cDNA of rye cytoplasmic ribosomal protein ScRPS7 as a query probe, a highly homologous rice genomic contig was obtained from Huada rice genome database. The full-length cDNA sequence of rice cytoplasmic ribosomal protein S7 was assembled by informatics based on the contig. Furthermore, with the two primers designed according to this assembled cDNA, the full-length cDNA of rice ribosomal protein was cloned by RT-PCR and named as OsRPS7. The cDNA was 919bp in length and contained a complete Open Reading Frame (ORF) of 576bp, encoding a protein of 192 amino acid residues. The deduced amino acids of OsRPS7 showed 88%,72% and 72% identity with those from Secale cereale, Arabidopsis thaliana and Brassica oleracea, respectively. The genome structure of OsRPS7 was analyzed, and its function was predicted in this paper.

Amino Acid Sequence↗

In search of immunogenic Helicobacter pylori proteins by screening of expression library.

BACKGROUND: Prevention of Helicobacter pylori infection may help to control related gastritis, peptic ulcer and cancer. Of the possible preventive measures, immunization was successfully employed in various animal studies. However, no immunization protocol has been accepted for humans. A better characterization of the immune response against the pathogen may be required before a human vaccine is developed. AIM: To identify bacterial proteins which induce an immune response in infected humans or H. pylori-immunized rabbits. METHODS: An expression library of H. pylori genes was screened with sera from infected humans and from immunized rabbits. Positive clones were partially sequenced and identified on the basis of a homology search of a H. pylori genome database. Encoded proteins were expressed directly from positive clones and analyzed by SDS-PAGE/Western blot techniques. RESULTS: 114 positive clones were isolated: 79 by screening with human sera and 35 by screening with rabbit sera. Western blot analysis demonstrated that selected clones encoded one or more strongly immunoreactive proteins. 64 clones selected with human sera had no counterparts among clones from screening with rabbit serum. 13 of these clones encoded a total of 21 unknown H. pylori proteins. 17 clones selected with rabbit sera were not immunostained with human sera. They represent 2 various regions of the H. pylori genome which encoded 3 bacterial proteins of unknown function. CONCLUSIONS: Screening of H. pylori expression library identified immunogenic proteins - potential vaccine antigens.

Animals↗

Identification, mapping, and genomic structure of a novel X-chromosomal human gene (SMPX) encoding a small muscular protein.

Reciprocal probing has been used to identify a cDNA clone (xh8H11) representing a gene preferentially expressed in striated muscle. The gene maps close to DXS7101 31.9 cM from the short arm telomere of the X-chromosome at Xp22.1. On searching expressed and genomic databases, 21 expressed sequence tags were found that allowed the assignment of a human extended consensus sequence of 887 bp, suggesting a completely expressed gene symbolized as SMPX. By using the human consensus sequence, the orthologous mouse Smpx and rat SMPX genes could be aligned and confirmed by complete sequencing of additional SMPX-related clones obtained by library screening. An open reading frame was identified encoding a peptide of 88-86 and 85 amino acids in human and rodents, respectively. The predicted peptide had no significant homologies to known structural elements. The human consensus cDNA sequence was used to define the genomic structure of the human SMPX that had been missed by a previous large scale sequencing approach. The gene consists of five exons (> or =172, 57, 84, 148, > or =422 bp) and four introns (3639, 10410, 6052, 31134 bp) comprising together 52.1 kb and is preferentially and abundantly expressed in heart and skeletal muscle. Thus, a novel human gene encoding a small muscular protein that maps to Xp22.1 (SMPX) has been identified and structurally characterized as a basis for further functional analysis.

Amino Acid Sequence↗

Use of the complete genome sequence information of Haemophilus influenzae strain Rd to investigate lipopolysaccharide biosynthesis.

The availability of the complete 1.83-megabase-pair sequence of the Haemophilus influenzae strain Rd genome has facilitated significant progress in investigating the biology of H.influenzae lipopolysaccharide (LPS), a major virulence determinant of this human pathogen. By searching the H. influenzae genomic database, with sequences of known LPS biosynthetic genes from other organisms, we identified and then cloned 25 candidate LPS genes. Construction of mutant strains and characterization of the LPS by reactivity with monoclonal antibodies, PAGE fractionation patterns and electrospray mass spectrometry comparative analysis have confirmed a potential role in LPS biosynthesis for the majority of these candidate genes. Virulence studies in the infant rat have allowed us to estimate the minimal LPS structure required for intravascular dissemination. This study is one of the first to demonstrate the rapidity, economy and completeness with which novel biological information can be accessed once the complete genome sequence of an organism is available.

Animals↗

Design and implementation of a qualitative simulation model of lambda phage infection.

MOTIVATION: Molecular biology databases hold a large number of empirical facts about many different aspects of biological entities. That data is static in the sense that one cannot ask a database 'What effect has protein A on gene B?' or 'Do gene A and gene B interact, and if so, how?'. Those questions require an explicit model of the target organism. Traditionally, biochemical systems are modelled using kinetics and differential equations in a quantitative simulator. For many biological processes however, detailed quantitative information is not available, only qualitative or fuzzy statements about the nature of interactions. RESULTS: We designed and implemented a qualitative simulation model of lambda phage growth control in Escherichia coli based on the existing simulation environment QSim. Qualitative reasoning can serve as the basis for automatic transformation of contents of genomic databases into interactive modelling systems that can reason about the relations and interactions of biological entities.

Algorithms↗

Histone acetylase GCN5 enters the nucleus via importin-alpha in protozoan parasite Toxoplasma gondii.

The histone acetyltransferase GCN5 acetylates nucleosomal histones to alter gene expression. How GCN5 gains entry into the nucleus of the cell has not been determined. We have mapped a six-amino acid motif (RKRVKR) that serves as a necessary and sufficient nuclear localization signal (NLS) for GCN5 in the protozoan pathogen Toxoplasma gondii (TgGCN5). Virtually nothing is known about nucleocytoplasmic transport in these parasites (phylum Apicomplexa), and this study marks the first demonstrated NLS delineated for members of the phylum. The TgGCN5 NLS has predictive value because it successfully identifies other nuclear proteins in three different apicomplexan genomic databases. Given the basic composition of the T. gondii NLS, we hypothesized that TgGCN5 physically interacts with importin-alpha, the main transport receptor in the importin/karyopherin nuclear import pathway. We cloned the importin-alpha gene from T. gondii (TgIMPalpha), which encodes a protein of 545 amino acids that possesses an importin-beta-binding domain and armadillo/beta-catenin-like repeats. In vitro co-immunoprecipitation experiments confirm that TgIMPalpha directly interacts with TgGCN5, but this interaction is abolished if the TgGCN5 NLS is deleted. Taken together, these data argue that TgGCN5 gains access to the parasite nucleus by interacting with TgIMPalpha. Bioinformatics analysis of the T. gondii genome reveals that other components of the importin pathway are present in the organism. This study demonstrates the utility of T. gondii as a model for the study of nucleocytoplasmic trafficking in early eukaryotic cells.

Acetyltransferases↗

G proteins, chemosensory perception, and the C. elegans genome project: An attractive story.

Heterotrimeric G proteins, consisting of alpha, beta, and gamma subunits, couple ligand-bound seven transmembrane domain receptors to the regulation of effector proteins and production of intracellular second messengers. G protein signaling mediates the perception of environmental cues in all higher eukaryotic organisms, including yeast, Dictyostelium, plants, and animals. The nematode Caenorhabditis elegans is the first animal to have complete descriptions of its cellular anatomy, cell lineage, neuronal wiring diagram, and genomic sequence. In a recent paper, Jansen et al. used sequence searches of the C. elegans genome database to identify all heterotrimeric G protein genes (20 Galpha, 2 Gbeta, 2 Ggamma). C. elegans encodes one ortholog of each of the four Galpha classes found in metazoans and 16 new Galpha genes. The orthologous genes are widely expressed, whereas 14 of the divergent Galpha genes are almost exclusively expressed in sensory neurons where they may regulate perception and chemotaxis.

Animals↗

EST-based gene discovery in pig: virtual expression patterns and comparative mapping to human.

A molecular understanding of porcine reproduction is of biological interest and economic importance. Our Midwest Consortium has produced cDNA libraries containing the majority of genes expressed in major female reproductive tissues, and we have deposited into public databases 21,499 expressed sequence tag (EST) gene sequences from the 3' end of clones from these libraries. These sequences represent 10,574 different genes, based on sequence comparison among these data, and comparison with existing porcine ESTs and genes indicate as many as 4652 of these EST clusters are novel. In silico analysis identified sequences that are expressed in specific pig tissues or organs and confirmed the broad expression in pig for many genes ubiquitously expressed in human tissues. Furthermore, we have developed computer software to identify sequence similarity of these pig genes with their human counterparts, and to extract the mapping information of these human homologues from genome databases. We demonstrate the utility of this software for comparative mapping by localizing 61 genes on the porcine physical map for Chromosomes (Chrs) 5, 10, and 14.

Algorithms↗

Characterization of a multisubunit transcription factor complex essential for spliced-leader RNA gene transcription in Trypanosoma brucei.

In the unicellular human parasites Trypanosoma brucei, Trypanosoma cruzi, and Leishmania spp., the spliced-leader (SL) RNA is a key molecule in gene expression donating its 5'-terminal region in SL addition trans splicing of nuclear pre-mRNA. While there is no evidence that this process exists in mammals, it is obligatory in mRNA maturation of trypanosomatid parasites. Hence, throughout their life cycle, these organisms crucially depend on high levels of SL RNA synthesis. As putative SL RNA gene transcription factors, a partially characterized small nuclear RNA-activating protein complex (SNAP(c)) and the TATA-binding protein related factor 4 (TRF4) have been identified thus far. Here, by tagging TRF4 with a novel epitope combination termed PTP, we tandem affinity purified from crude T. brucei extracts a stable and transcriptionally active complex of six proteins. Besides TRF4 these were identified as extremely divergent subunits of SNAP(c) and of transcription factor IIA (TFIIA). The latter finding was unexpected since genome databases of trypanosomatid parasites appeared to lack general class II transcription factors. As we demonstrate, the TRF4/SNAP(c)/TFIIA complex binds specifically to the SL RNA gene promoter upstream sequence element and is absolutely essential for SL RNA gene transcription in vitro.

5' Untranslated Regions↗

Single nucleotide polymorphisms in cytochrome P450 genes from barley.

Plant cytochrome P450s are known to be essential in a number of economically important pathways of plant metabolism but there are also many P450s of unknown function accumulating in expressed sequence tag (EST) and genomic databases. To detect trait associations that could assist in the assignment of gene function and provide markers for breeders selecting for commercially important traits, detection of polymorphisms in identified P450 genes is desirable. Polymorphisms in EST sequences provide so-called perfect markers for the associated genes. The International Triticeae EST Cooperative data base of 24,344 ESTs was searched for sequences exhibiting homology to P450 genes representing the nine known clans of plant P450s. Seventy five P450 ESTs were identified of which 24 had best matches in Genbank to P450 genes of known function and 51 to P450s of unknown function. Sequence information from PCR products amplified from the genomic template DNA of 11 barley varieties was obtained using primers designed from six barley P450 ESTs and one durum wheat P450 EST. Single nucleotide polymorphisms (SNPs) between barley varieties were identified using five of the seven PCR products. A maximum of five SNPs and three haplotypes among the 11 barley lines were detected in products from any one primer pair. SNPs in three PCR products led to changes between barley varieties in at least one restriction site enabling genotyping and mapping without the expense of a specialist SNP detection system. The overall frequency of SNPs across the 11 barley varieties was 1 every 131 bases.

Base Sequence↗

Molecular control of the oocyte to embryo transition.

The elucidation of the molecular control of the initiation of mammalian embryogenesis is possible now that the transcriptomes of the full-grown oocyte and two-cell stage embryo have been prepared and analysed. Functional annotation of the transcriptomes using gene ontology vocabularies, allows comparison of the oocyte and two-cell stage embryo between themselves, and with all known mouse genes in the Mouse Genome Database. Using this methodology one can outline the general distinguishing features of the oocyte and the two-cell stage embryo. This, when combined with oocyte-specific targeted deletion of genes, allows us to dissect the molecular networks at play as the differentiated oocyte and sperm transit into blastomeres with unlimited developmental potential.

Animals↗

Bacterial degradation of xenobiotic compounds: evolution and distribution of novel enzyme activities.

Bacterial dehalogenases catalyse the cleavage of carbon-halogen bonds, which is a key step in aerobic mineralization pathways of many halogenated compounds that occur as environmental pollutants. There is a broad range of dehalogenases, which can be classified in different protein superfamilies and have fundamentally different catalytic mechanisms. Identical dehalogenases have repeatedly been detected in organisms that were isolated at different geographical locations, indicating that only a restricted number of sequences are used for a certain dehalogenation reaction in organohalogen-utilizing organisms. At the same time, massive random sequencing of environmental DNA, and microbial genome sequencing projects have shown that there is a large diversity of dehalogenase sequences that is not employed by known catabolic pathways. The corresponding proteins may have novel functions and selectivities that could be valuable for biotransformations in the future. Apparently, traditional enrichment and metagenome approaches explore different segments of sequence space. This is also observed with alkane hydroxylases, a category of proteins that can be detected on basis of conserved sequence motifs and for which a large number of sequences has been found in isolated bacterial cultures and genomic databases. It is likely that ongoing genetic adaptation, with the recruitment of silent sequences into functional catabolic routes and evolution of substrate range by mutations in structural genes, will further enhance the catabolic potential of bacteria toward synthetic organohalogens and ultimately contribute to cleansing the environment of these toxic and recalcitrant chemicals.

Amino Acid Sequence↗

Nuclear heat shock response and novel nuclear domain 10 reorganization in respiratory syncytial virus-infected a549 cells identified by high-resolution two-dimensional gel electrophoresis.

The pneumovirus respiratory syncytial virus (RSV) is a leading cause of epidemic respiratory tract infection. Upon entry, RSV replicates in the epithelial cytoplasm, initiating compensatory changes in cellular gene expression. In this study, we have investigated RSV-induced changes in the nuclear proteome of A549 alveolar type II-like epithelial cells by high-resolution two-dimensional gel electrophoresis (2DE). Replicate 2D gels from uninfected and RSV-infected nuclei were compared for changes in protein expression. We identified 24 different proteins by peptide mass fingerprinting after matrix-assisted laser desorption ionization-time of flight mass spectrometry (MS), whose average normalized spot intensity was statistically significant and differed by +/-2-fold. Notable among the proteins identified were the cytoskeletal cytokeratins, RNA helicases, oxidant-antioxidant enzymes, the TAR DNA binding protein (a protein that associates with nuclear domain 10 [ND10] structures), and heat shock protein 70- and 60-kDa isoforms (Hsp70 and Hsp60, respectively). The identification of Hsp70 was also validated by liquid chromatography quadropole-TOF tandem MS (LC-MS/MS). Separate experiments using immunofluorescence microscopy revealed that RSV induced cytoplasmic Hsp70 aggregation and nuclear accumulation. Data mining of a genomic database showed that RSV replication induced coordinate changes in Hsp family proteins, including the 70, 70-2, 90, 40, and 40-3 isoforms. Because the TAR DNA binding protein associates with ND10s, we examined the effect of RSV infection on ND10 organization. RSV induced a striking dissolution of ND10 structures with redistribution of the component promyelocytic leukemia (PML) and speckled 100-kDa (Sp100) proteins into the cytoplasm, as well as inducing their synthesis. Our findings suggest that cytoplasmic RSV replication induces a nuclear heat shock response, causes ND10 disruption, and redistributes PML and Sp100 to the cytoplasm. Thus, a high-resolution proteomics approach, combined with immunofluorescence localization and coupled with genomic response data, yielded unexpected novel insights into compensatory nuclear responses to RSV infection.

Cell Nucleus↗

A surrogate-based approach for post-genomic partner identification.

BACKGROUND: Modern drug discovery is concerned with identification and validation of novel protein targets from among the 30,000 genes or more postulated to be present in the human genome. While protein-protein interactions may be central to many disease indications, it has been difficult to identify new chemical entities capable of regulating these interactions as either agonists or antagonists. RESULTS: In this paper, we show that peptide complements (or surrogates) derived from highly diverse random phage display libraries can be used for the identification of the expected natural biological partners for protein and non-protein targets. Our examples include surrogates isolated against both an extracellular secreted protein (TNFbeta) and intracellular disease related mRNAs. In each case, surrogates binding to these targets were obtained and found to contain partner information embedded in their amino acid sequences. Furthermore, this information was able to identify the correct biological partners from large human genome databases by rapid and integrated computer based searches. CONCLUSIONS: Modified versions of these surrogates should provide agents capable of modifying the activity of these targets and enable one to study their involvement in specific biological processes as a means of target validation for downstream drug discovery.

Computational Biology↗