Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Use of a Mycobacterium tuberculosis H37Rv bacterial artificial chromosome library for genome mapping, sequencing, and comparative genomics.

The bacterial artificial chromosome (BAC) cloning system is capable of stably propagating large, complex DNA inserts in Escherichia coli. As part of the Mycobacterium tuberculosis H37Rv genome sequencing project, a BAC library was constructed in the pBeloBAC11 vector and used for genome mapping, confirmation of sequence assembly, and sequencing. The library contains about 5,000 BAC clones, with inserts ranging in size from 25 to 104 kb, representing theoretically a 70-fold coverage of the M. tuberculosis genome (4.4 Mb). A total of 840 sequences from the T7 and SP6 termini of 420 BACs were determined and compared to those of a partial genomic database. These sequences showed excellent correlation between the estimated sizes and positions of the BAC clones and the sizes and positions of previously sequenced cosmids and the resulting contigs. Many BAC clones represent linking clones between sequenced cosmids, allowing full coverage of the H37Rv chromosome, and they are now being shotgun sequenced in the framework of the H37Rv sequencing project. Also, no chimeric, deleted, or rearranged BAC clones were detected, which was of major importance for the correct mapping and assembly of the H37Rv sequence. The minimal overlapping set contains 68 unique BAC clones and spans the whole H37Rv chromosome with the exception of a single gap of approximately 150 kb. As a postgenomic application, the canonical BAC set was used in a comparative study to reveal chromosomal polymorphisms between M. tuberculosis, M. bovis, and M. bovis BCG Pasteur, and a novel 12.7-kb segment present in M. tuberculosis but absent from M. bovis and M. bovis BCG was characterized. This region contains a set of genes whose products show low similarity to proteins involved in polysaccharide biosynthesis. The H37Rv BAC library therefore provides us with a powerful tool both for the generation and confirmation of sequence data as well as for comparative genomics and other postgenomic applications. It represents a major resource for present and future M. tuberculosis research projects.

Chromosome Mapping↗

Is retinoic acid genetic machinery a chordate innovation?

Development of many chordate features depends on retinoic acid (RA). Because the action of RA during development seems to be restricted to chordates, it had been previously proposed that the "invention" of RA genetic machinery, including RA-binding nuclear hormone receptors (Rars), and the RA-synthesizing and RA-degrading enzymes Aldh1a (Raldh) and Cyp26, respectively, was an important step for the origin of developmental mechanisms leading to the chordate body plan. We tested this hypothesis by conducting an exhaustive survey of the RA machinery in genomic databases for twelve deuterostomes. We reconstructed the evolution of these genes in deuterostomes and showed for the first time that RA genetic machinery--that is Aldh1a, Cyp26, and Rar orthologs--is present in nonchordate deuterostomes. This finding implies that RA genetic machinery was already present during early deuterostome evolution, and therefore, is not a chordate innovation. This new evolutionary viewpoint argues against the hypothesis that the acquisition of gene families underlying RA metabolism and signaling was a key event for the origin of chordates. We propose a new hypothesis in which lineage-specific duplication and loss of RA machinery genes could be related to the morphological radiation of deuterostomes.

Aldehyde Oxidoreductases↗

Human skeletal muscle triadin: gene organization and cloning of the major isoform, Trisk 51.

We obtained the gene organization of human triadin gene by aligning the DNA coding sequence of human 95-kDa triadin (Trisk 95) with human genomic database. We identified a novel human triadin isoform, a potential human homologue of rat Trisk 51. We show that both isoforms of triadin, Trisk 51 and Trisk 95, are alternative splice variants of the same gene. We demonstrated experimentally the existence of this Trisk 51 transcript in human skeletal muscle and cloned its full length cDNA. We further demonstrated that the protein encoded by this transcript is expressed in the human skeletal muscle. In addition, unlike other species, Trisk 51 is the major triadin isoform expressed in human skeletal muscle, whereas Trisk 95 is below the detection level in the two types of muscles tested.

Amino Acid Sequence↗

DNM1DN: a new class of paralogous genomic segments (duplicons) with highly conserved copies on chromosomes Y and 15.

Screening a testis cDNA selection library for Y-linked genes yielded 79 cDNAs. Of these, 9 matched the 3' region of the dynamin 1 gene (DNM1) on chromosome 9q34 with >90% identity. Fluoresence in situ hybridisation and PCR amplification were used to localise a large number of DNM1-like sequences to human chromosomes 15 and Y. PCR amplification of overlapping Y-linked YACs allowed a more accurate mapping of the Y-linked DNM1-like cDNAs to a euchromatic locus in close proximity to heterochromatin at Yq11.23. A search of the genome database identified 64 highly homologous copies of the DNM1 fragment. Most of these copies were localised to chromosomes 15 and Y, but others mapped to chromosomes 5, 8, 10, 12, 19 and 22. These sequences exhibit all the major features of a duplicon and have been designated DNM1DN (DNM1 duplicon). Evolutionary studies using fluorescence in situ hybridisation indicate that transposition of the DNM1DN sequence to chromosome 15 took place earlier in primate evolution than the transposition to the Y chromosome. The translocation to the Y took place at a time following the divergence of a common ancestor from gorilla, approximately 4-7 million years ago.

Animals↗

Isolation and functional analysis of human HMBOX1, a homeobox containing protein with transcriptional repressor activity.

We have identified and isolated a novel human gene, HMBOX1 (homeobox containing 1) from a pancreatic cDNA library. Human HMBOX1 is widely expressed in 18 tissues, and it is highly expressed in pancreas. According to the genome database, HMBOX1 is located at the boundary of 8p12.3 and 8p21.1. HMBOX1 proteins are highly conserved in human, mouse, rat, chicken and Xenopus laevis. A phylogenetic tree shows that HMBOX1 may represent a distinct group in HNF (Hepatocyte Nuclear Factor) transcriptional factors. Functional HMBOX1::EGFP (enhanced green fluorescent protein) fusion protein revealed that HMBOX1 accumulated more in cytoplasm than in nucleus. Co-transfection of HEK-293T cells with pM-HMBOX1 plasmid and reporter plasmid pGAL4(5)tkLUC indicates that HMBOX1 is a transcription repressor. In situ hybridization on paraffin sections of mouse tissues demonstrated that Hmbox1 is widely expressed in pancreas and the expression of this gene can also be detected in pallium, hippocampus and hypothalamus.

Amino Acid Sequence↗

Identification of CD55 as a downstream factor of EP4 receptor signaling in colorectal cancer cells.

Prostaglandin E2 (PGE2) signaling through the E-type prostanoid 4 (EP4) receptor has been implicated in the pathophysiology of colorectal cancer (CRC). We herein identified decay-accelerating factor, also known as CD55, as a novel CRC-associated downstream factor of the EP4 receptor. The integration of transcriptomic profiling of PGE2-stimulated HCA-7 human colon cancer cells with analyses of cancer genomic databases predicted CD55 as a potential EP4 receptor-regulated target. Inhibitor-based experiments showed the induction of CD55 after a PGE2 stimulation required the EP4 receptor and Gi protein in HCA-7 cells, whereas protein kinase A signaling was dispensable. In combination with a toxicogenomic database analysis, p38 mitogen-activated protein kinase (MAPK) was identified as the predominant effector connecting the EP4 receptor to CD55 upregulation. A single-cell RNA-seq re-analysis of human CRC tissues revealed CD55 upregulation and p38 MAPK-related gene set enrichment in epithelial cells expressing the EP4 receptor, suggesting that this induction mechanism may operate in a subset of epithelial cells in clinical specimens. Collectively, these results delineate a PGE2/EP4 receptor/Gi protein/p38 MAPK signaling axis that induces CD55 expression in HCA-7 cells and epithelial tumor cells, provide new mechanistic clues for understanding the regulation of complement regulatory molecule CD55 expression by prostaglandin signaling.

Humans↗

Evolution and expression of chimeric POTE-actin genes in the human genome.

We previously described a primate-specific gene family, POTE, that is expressed in many cancers but in a limited number of normal organs. The 13 POTE genes are dispersed among eight different chromosomes and evolved by duplications and remodeling of the human genome from an ancestral gene, ANKRD26. Based on sequence similarity, the POTE gene family members can be divided into three groups. By genome database searches, we identified an actin retroposon insertion at the carboxyl terminus of one of the ancestral POTE paralogs. By Northern blot analysis, we identified the expected 7.5-kb POTE-actin chimeric transcript in a breast cancer cell line. The protein encoded by the POTE-actin transcript is predicted to be 120 kDa in size. Using anti-POTE mAbs that recognize the amino-terminal portion of the POTE protein, we detected the 120-kDa POTE-actin fusion protein in breast cancer cell lines known to express the fusion transcript. These data demonstrate that insertion of a retroposon produced an altered functional POTE gene. This example indicates that new functional human genes can evolve by insertion of retroposons.

Actins↗

Sequence-based heuristics for faster annotation of non-coding RNA families.

MOTIVATION: Non-coding RNAs (ncRNAs) are functional RNA molecules that do not code for proteins. Covariance Models (CMs) are a useful statistical tool to find new members of an ncRNA gene family in a large genome database, using both sequence and, importantly, RNA secondary structure information. Unfortunately, CM searches are extremely slow. Previously, we created rigorous filters, which provably sacrifice none of a CM's accuracy, while making searches significantly faster for virtually all ncRNA families. However, these rigorous filters make searches slower than heuristics could be. RESULTS: In this paper we introduce profile HMM-based heuristic filters. We show that their accuracy is usually superior to heuristics based on BLAST. Moreover, we compared our heuristics with those used in tRNAscan-SE, whose heuristics incorporate a significant amount of work specific to tRNAs, where our heuristics are generic to any ncRNA. Performance was roughly comparable, so we expect that our heuristics provide a high-quality solution that--unlike family-specific solutions--can scale to hundreds of ncRNA families. AVAILABILITY: The source code is available under GNU Public License at the supplementary web site.

Algorithms↗

A novel homeobox gene overexpressed in thyroid carcinoma.

Serial analysis of gene expression (SAGE) was applied to compare expression profiles of normal thyroid tissue and papillary thyroid carcinoma (PTC). A SAGE tag corresponding to the partial cDNA for the small protein 31 (SMAP31) is upregulated approximately 13-fold in papillary thyroid cancer (PTC) and was selected for further research. BLAST-searching the human genome database reveals that the SMAP31 gene is located on chromosome 4q11-12 and contains 6 exons. Alternative splicing results in seven transcripts encoding 2 possible open reading frames (ORF) of 73 and 95 amino acids. Database searching in GenBank's dbEST shows that SMAP31 transcripts are expressed mainly in brain, heart, gingiva, and lung tissue. Thyroid tissue contains three transcripts caused by alternatively splicing in the 5' untranslated region (UTR), which all encode an identical ORF of 73 amino acids. Homology search shows that this protein contains a homeobox domain. Thyroid and/or thyroid carcinoma-specific expression of SMAP31 is studied using Northern blot and reverse transcriptase-polymerase chain reaction (RT-PCR) on a multiple tissue panel. RT-PCR experiments on a cDNA panel containing samples from different normal and tumor tissues shows expression of SMAP31 mRNA in brain, placenta, lung, heart, thyroid and thyroid carcinoma. SMAP31 expression is elevated in 4 of 6 PTC tumor samples compared to 4 normal thyroid controls.

Alternative Splicing↗

Characterization of the pufferfish Takifugu rubripes apolipoprotein multigene family.

We have characterized the apolipoprotein multigene family of the pufferfish Takifugu rubripes. The pufferfish mainly contains 28-kDa, 27-kDa, and 14-kDa apolipoproteins in its plasma and was designated apo-28 kDa, apo-27 kDa, and apo-14 kDa, respectively. N-terminal amino acid sequencing revealed that pufferfish apo-28 kDa and apo-27 kDa have an identical amino acid sequence except an additional propeptide in the former; and both are homologues of apoA-I from other animals. The sequence of pufferfish apo-14 kDa is homologous to that of eel apo-14 kDa previously reported, both being apparently specific to fish. In silico screening, using the publicly available Fugu genome database confirmed the pufferfish apoA-I and apo-14 kDa genes. The database further contained the genes encoding four types of apoA-IV, one apoC-II and two types of apoE. Thus, pufferfish contains nine genes encoding apolipoprotein multigene family. Two apoA-IV and one apoE genes were tandemly arrayed and located on one scaffold. Thus two sets of these genes formed two gene clusters. The apoC-II and apo-14 kDa genes are also located on a single scaffold. apoA-I and apo-14 kDa gene transcripts were mainly expressed in liver and less abundantly in brain. The transcripts of the former gene were also observed in intestine. In contrast, the transcripts encoding four apoA-IVs, one apoC-II, and two apoEs were mainly expressed in intestine. These structural details of pufferfish apolipoproteins and tissue distribution of their gene transcripts provide a novel evidence for better understanding of evolutionary relationships of apolipoprotein multigene family.

Amino Acid Sequence↗

Fe-hydrogenase maturases in the hydrogenosomes of Trichomonas vaginalis.

Assembly of active Fe-hydrogenase in the chloroplasts of the green alga Chlamydomonas reinhardtii requires auxiliary maturases, the S-adenosylmethionine-dependent enzymes HydG and HydE and the GTPase HydF. Genes encoding homologous maturases had been found in the genomes of all eubacteria that contain Fe-hydrogenase genes but not yet in any other eukaryote. By means of proteomic analysis, we identified a homologue of HydG in the hydrogenosomes, mitochondrion-related organelles that produce hydrogen under anaerobiosis by the activity of Fe-hydrogenase, in the pathogenic protist Trichomonas vaginalis. Genes encoding two other components of the Hyd system, HydE and HydF, were found in the T. vaginalis genome database. Overexpression of HydG, HydE, and HydF in trichomonads showed that all three proteins are specifically targeted to the hydrogenosomes, the site of Fe-hydrogenase maturation. The results of Neighbor-Net analyses of sequence similarities are consistent with a common eubacterial ancestor of HydG, HydE, and HydF in T. vaginalis and C. reinhardtii, supporting a monophyletic origin of Fe-hydrogenase maturases in the two eukaryotes. Although Fe-hydrogenases exist in only a few eukaryotes, related Narf proteins with different cellular functions are widely distributed. Thus, we propose that the acquisition of Fe-hydrogenases, together with Hyd maturases, occurred once in eukaryotic evolution, followed by the appearance of Narf through gene duplication of the Fe-hydrogenase gene and subsequent loss of the Hyd proteins in eukaryotes in which Fe-hydrogenase function was lost.

Amino Acid Motifs↗

Characterization of a Candida albicans gene encoding a putative transcriptional factor required for cell wall integrity.

After screening a Candida albicans genome database the product of an open reading frame (ORF) (CA2880) with 49% homology to the product of Saccharomyces cerevisiae YPL133c, a putative transcriptional factor, was identified. The disruption of the C. albicans gene leads to a major sensitivity to calcofluor white and Congo red, a minor sensitivity to sodium dodecyl sulfate, a major resistance to zymolyase, and an alteration of the chemical composition of the cell wall. For these reasons we called it CaCWT1 (for C. albicans cell wall transcription factor). CaCwt1p contains a putative Zn(II) Cys(6) DNA binding domain characteristic of some transcriptional factors and a PAS domain. The CaCWT1 gene is more expressed in stationary phase cells than in cells growing exponentially. To our knowledge, this is the first Zn(II) Cys(6) transcriptional factor-encoding gene implicated in the cell wall architecture.

Amino Acid Sequence↗

The 'permeome' of the malaria parasite: an overview of the membrane transport proteins of Plasmodium falciparum.

BACKGROUND: The uptake of nutrients, expulsion of metabolic wastes and maintenance of ion homeostasis by the intraerythrocytic malaria parasite is mediated by membrane transport proteins. Proteins of this type are also implicated in the phenomenon of antimalarial drug resistance. However, the initial annotation of the genome of the human malaria parasite Plasmodium falciparum identified only a limited number of transporters, and no channels. In this study we have used a combination of bioinformatic approaches to identify and attribute putative functions to transporters and channels encoded by the malaria parasite, as well as comparing expression patterns for a subset of these. RESULTS: A computer program that searches a genome database on the basis of the hydropathy plots of the corresponding proteins was used to identify more than 100 transport proteins encoded by P. falciparum. These include all the transporters previously annotated as such, as well as a similar number of candidate transport proteins that had escaped detection. Detailed sequence analysis enabled the assignment of putative substrate specificities and/or transport mechanisms to all those putative transport proteins previously without. The newly-identified transport proteins include candidate transporters for a range of organic and inorganic nutrients (including sugars, amino acids, nucleosides and vitamins), and several putative ion channels. The stage-dependent expression of RNAs for 34 candidate transport proteins of particular interest are compared. CONCLUSION: The malaria parasite possesses substantially more membrane transport proteins than was originally thought, and the analyses presented here provide a range of novel insights into the physiology of this important human pathogen.

Amino Acid Sequence↗

Promoter characterization and genomic organization of the human breast cancer resistance protein (ATP-binding cassette transporter G2) gene.

The breast cancer resistance protein (BCRP) gene, formally known as ATP-binding cassette transporter G2 (ABCG2) gene, encodes an ABC half transporter that causes resistance to certain cancer chemotherapeutic drugs when transfected and expressed in drug sensitive cancer cells. Here we report the organization of the BCRP gene, and the initial characterization of the BCRP promoter. We identified the genomic sequence of BCRP and its promoter by screening a human genomic lambda phage library, as well as a BAC library, and by searching the human genome database. The BCRP gene spans over 66 kb and consists of 16 exons and 15 introns. The exons range in size from 60 to 532 bp. The translational start site is found in the second exon. The first exon contains the majority of the 5' UTR. Promoter activity was characterized by a luciferase reporter assay using transient transfection of the human breast cancer cell line MCF7, and the human choriocarcinoma cell lines JAR, BeWo and JEG-3, which we find to have high endogenous expression of BCRP. The BCRP gene is transcribed by a TATA-less promoter with several putative Sp1 sites, which are downstream from a putative CpG island. The sequence 312 bp directly upstream from the BCRP transcriptional start site conferred basal promoter activity. The 5' region upstream of the basal promoter is characterized by both positive and negative regulatory domains.

ATP Binding Cassette Transporter, Subfamily G, Mem↗

The mouse olfactory receptor gene family.

In mammals, odor detection in the nose is mediated by a diverse family of olfactory receptors (ORs), which are used combinatorially to detect different odorants and encode their identities. The OR family can be divided into subfamilies whose members are highly related and are likely to recognize structurally related odorants. To gain further insight into the mechanisms underlying odor detection, we analyzed the mouse OR gene family. Exhaustive searches of a mouse genome database identified 913 intact OR genes and 296 OR pseudogenes. These genes were localized to 51 different loci on 17 chromosomes. Sequence comparisons showed that the mouse OR family contains 241 subfamilies. Subfamily sizes vary extensively, suggesting that some classes of odorants may be more easily detected or discriminated than others. Determination of subfamilies that contain ORs with identified ligands allowed tentative functional predictions for 19 subfamilies. Analysis of the chromosomal locations of members of each subfamily showed that many OR gene loci encode only one or a few subfamilies. Furthermore, most subfamilies are encoded by a single locus, suggesting that different loci may encode receptors for different types of odorant structural features. Comparison of human and mouse OR subfamilies showed that the two species have many, but not all, subfamilies in common. However, mouse subfamilies are usually larger than their human counterparts. This finding suggests that humans and mice recognize many of the same odorant structural motifs, but mice may be superior in odor sensitivity and discrimination.

Animals↗

Software agents in molecular computational biology.

Progress made in applying agent systems to molecular computational biology is reviewed and strategies by which to exploit agent technology to greater advantage are investigated. Communities of software agents could play an important role in helping genome scientists design reagents for future research. The advent of genome sequencing in cattle and swine increases the complexity of data analysis required to conduct research in livestock genomics. Databases are always expanding and semantic differences among data are common. Agent platforms have been developed to deal with generic issues such as agent communication, life cycle management and advertisement of services (white and yellow pages). This frees computational biologists from the drudgery of having to re-invent the wheel on these common chores, giving them more time to focus on biology and bioinformatics. Agent platforms that comply with the Foundation for Intelligent Physical Agents (FIPA) standards are able to interoperate. In other words, agents developed on different platforms can communicate and cooperate with one another if domain-specific higher-level communication protocol details are agreed upon between different agent developers. Many software agent platforms are peer-to-peer, which means that even if some of the agents and data repositories are temporarily unavailable, a subset of the goals of the system can still be met. Past use of software agents in bioinformatics indicates that an agent approach should prove fruitful. Examination of current problems in bioinformatics indicates that existing agent platforms should be adaptable to novel situations.

Algorithms↗

Expression of beta-expansins is correlated with internodal elongation in deepwater rice.

Fourteen putative rice (Oryza sativa) beta-expansin genes, Os-EXPB1 through Os-EXPB14, were identified in the expressed sequence tag and genomic databases. The DNA and deduced amino acid sequences are highly conserved in all 14 beta-expansins. They have a series of conserved C (cysteine) residues in the N-terminal half of the protein, an HFD (histidine-phenylalanine-aspartate) motif in the central region, and a series of W (tryptophan) residues near the carboxyl terminus. Five beta-expansin genes are expressed in deepwater rice internodes, with especially high transcript levels in the growing region. Expression of four beta-expansin genes in the internode was induced by treatment with gibberellin and by wounding. The wound response resulted from excising stem sections or from piercing pinholes into the stem of intact plants. The level of wound-induced beta-expansin transcripts declined rapidly 5 h after cutting of stem sections. We conclude that the expression of beta-expansin genes is correlated with rapid elongation of deepwater rice internodes, it is induced by gibberellin and wounding, and wound-induced beta-expansin mRNA appears to turn over rapidly.

Adaptation, Physiological↗

Analysis of lung tumorigenesis in chimeric mice indicates the Pulmonary adenoma resistance 2 (Par2) locus to operate in the tumor-initiation stage in a cell-autonomous manner: detection of polymorphisms in the Poli gene as a candidate for Par2.

The Pulmonary adenoma resistance 2 (Par2) locus of the BALB/cByJ mouse, located within 0.5 cM of chromosome 18, is responsible for reducing the mean multiplicity of urethane-induced lung tumors relative to those in C57BL/6J, A/J and C3H/HeJ mice. Thus, BALB/B6-Par2 congenic strain genetically identical to BALB/cByJ except carrying C57BL/6J Par2 alleles develops seven times more tumors than BALB/cByJ. To gain clues for identification of Par2 candidate genes, we analysed lung tumorigenesis in BALB/cByJ<-->BALB.B6-Par2 chimeric animals. Of 100 tumors induced by urethane in 16 chimeras, 82 originated from BALB.B6-Par2 cells, indicating the Par2 phenotype to be cell-autonomous. In addition, the BALB.B6-Par2- and BALB/cByJ-derived tumors were similar in mean size, implying that the phenotype is primarily expressed during initiation rather than in the promotion stage of carcinogenesis. Given these results, we surveyed a comprehensive mouse genome database and physically mapped Par2 within a 2.3 Mbp segment containing three known genes, Poli, Mbd2 and Dcc. Among those, the Poli seemed to be the most reasonable Par2 candidate, since it encodes an extremely error-prone DNA polymerase preferentially incorporating G or T opposite template T in vitro, reminiscent of the Kras2 activation because of an A to G or T point mutation within codon 61 with which most urethane-induced lung tumors are initiated. Indeed, our sequencing of Poli cDNAs from BALB/cByJ, C57BL/6J, A/J and C3H/HeJ lungs revealed 21 BALB/cByJ-specific single-nucleotide polymorphisms in the coding region accompanied by seven amino-acid substitutions and an elevated frequency of alternative splicing, while no polymorphisms associated with tumor susceptibility were found for either Mbd2 or Dcc. Notably, we obtained evidence that BALB/cByJ Par2 alleles may selectively decrease the frequency of Kras2-mutated tumors compared with C57BL/6J alleles. Consequently, the Poli is an intriguing Par2 candidate clearly deserving further evaluation.

Adenoma↗