Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Identification of novel tropomyosin 1 genes of pufferfish (Fugu rubripes) on genomic sequences and tissue distribution of their transcripts.

Fugu genome database enabled us to identify two novel tropomyosin 1 (TPM1) genes through in silico data mining and isolation of their corresponding cDNAs in vivo. The duplicate TPM1 genes in Japanese pufferfish Fugu rubripes suggest that additional an ancient segmental duplication or whole genome duplication occurred in fish lineage, which, like many other reported Fugu genes, showed reduction in genomic size in comparison with their human homologue. Computer analysis predicted that the coiled-coil probabilities, that were thought to be the most major function of TPM, were the same between the two TPM1 isoforms. We confirmed that the tissue expression profiles of the two TPM1 genes differed from each other, which implied that changes in expression pattern could fix duplicated TPM1 genes although the two TPM1 isoforms appear to have similar function.

Animals↗

Extensive sequence homology between the mycobacterium leprae LSR (12 kDa) antigen and its Mycobacterium tuberculosis counterpart.

The Mycobacterium leprae LSR (12 kDa) protein antigen has been reported to mimic whole cell M. leprae in T cell responses across the leprosy spectrum. In addition, B cell responses to specific sequences within the LSR antigen have been shown to be associated with immunopathological responses in leprosy patients with erythema nodosum leprosum. We have in the present study applied the M. leprae LSR DNA sequence as query to search for the presence of homologous genes within the recently completed Mycobacterium tuberculosis genome database (Sanger Centre, UK). By using the BLASTN search tool, a homologous M. tuberculosis open reading frame (336 bp), encoding a protein antigen of 12.1 kDa, was identified within the cosmid MTCY07H7B.25. The gene is designated Rv3597c within the M. tuberculosis H37Rv genome. Sequence alignment revealed 93% identity between the M. leprae and M. tuberculosis antigens at the amino acid sequence level. The finding that some B and T cell epitopes were localized to regions with amino acid substitutions may account for the putative differential responsiveness to this antigen in tuberculosis and leprosy.

Amino Acid Sequence↗

Application of physiological genomics to the study of hearing disorders.

Although the biophysical principles of how the ear operates are reasonably well understood, little is known about the specific genes that confer normal function to the inner ear. Nevertheless, the recent implementation of genomic tools has led to extraordinary progress in the identification of mutated genes that cause non-syndromic and syndromic forms of deafness. Part of this success is directly related to the sequencing of the human and mouse genomes and improved gene annotation methods. This review discusses how physiological genomic tools, such as genomic databases, expressed sequence tag databases and DNA arrays have been applied to find candidate genes for important molecular processes in the inner ear. It also illustrates, using the discovery of genes encoding essential components of cochlear K+ homeostasis as an example, how the combination of physiological genomic tools with physiological and morphological information has led to an in-depth understanding of cochlear ion homeostasis. Finally, it discusses how the use of applied genomic tools, such as gene arrays, will further advance our knowledge of how the inner ear works, develops, ages and regenerates.

Cochlea↗

Integrating genomic data to predict transcription factor binding.

Transcription factor binding sites (TFBS) in gene promoter regions are often predicted by using position specific scoring matrices (PSSMs), which summarize sequence patterns of experimentally determined TF binding sites. Although PSSMs are more reliable than simple consensus string matching in predicting a true binding site, they generally result in high numbers of false positive hits. This study attempts to reduce the number of false positive matches and generate new predictions by integrating various types of genomic data by two methods: a Bayesian allocation procedure, and support vector machine classification. Several methods will be explored to strengthen the prediction of a true TFBS in the Saccharomyces cerevisiae genome: binding site degeneracy, binding site conservation, phylogenetic profiling, TF binding site clustering, gene expression profiles, GO functional annotation, and k-mer counts in promoter regions. Binding site degeneracy (or redundancy) refers to the number of times a particular transcription factor's binding motif is discovered in the upstream region of a gene. Phylogenetic conservation takes into account the number of orthologous upstream regions in other genomes that contain a particular binding site. Phylogenetic profiling refers to the presence or absence of a gene across a large set of genomes. Binding site clusters are statistically significant clusters of TF binding sites detected by the algorithm ClusterBuster. Gene expression takes into account the idea that when the gene expression profiles of a transcription factor and a potential target gene are correlated, then it is more likely that the gene is a genuine target. Also, genes with highly correlated expression profiles are often regulated by the same TF(s). The GO annotation data takes advantage of the idea that common transcription targets often have related function. Finally, the distribution of the counts of all k-mers of length 4, 5, and 6 in gene's promoter region were examined as means to predict TF binding. In each case the data are compared to known true positives taken from ChIP-chip data, Transfac, and the Saccharomyces Genome Database. First, degeneracy, conservation, expression, and binding site clusters were examined independently and in combination via Bayesian allocation. Then, binding sites were predicted with a support vector machine (SVM) using all methods alone and in combination. The SVM works best when all genomic data are combined, but can also identify which methods contribute the most to accurate classification. On average, a support vector machine can classify binding sites with high sensitivity and an accuracy of almost 80%.

Algorithms↗

6-Phosphofructo-2-kinase and fructose-2,6-bisphosphatase in Trypanosomatidae. Molecular characterization, database searches, modelling studies and evolutionary analysis.

Fructose 2,6-bisphosphate is a potent allosteric activator of trypanosomatid pyruvate kinase and thus represents an important regulator of energy metabolism in these protozoan parasites. A 6-phosphofructo-2-kinase, responsible for the synthesis of this regulator, was highly purified from the bloodstream form of Trypanosoma brucei and kinetically characterized. By searching trypanosomatid genome databases, four genes encoding proteins homologous to the mammalian bifunctional enzyme 6-phosphofructo-2-kinase/fructose-2,6-bisphosphatase (PFK-2/FBPase-2) were found for both T. brucei and the related parasite Leishmania major and four pairs in Trypanosoma cruzi. These genes were predicted to each encode a protein in which, at most, only a single domain would be active. Two of the T. brucei proteins showed most conservation in the PFK-2 domain, although one of them was predicted to be inactive due to substitution of residues responsible for ligating the catalytically essential divalent metal cation; the two other proteins were most conserved in the FBPase-2 domain. The two PFK-2-like proteins were expressed in Escherichia coli. Indeed, the first displayed PFK-2 activity with similar kinetic properties to that of the enzyme purified from T. brucei, whereas no activity was found for the second. Interestingly, several of the predicted trypanosomatid PFK-2/FBPase-2 proteins have long N-terminal extensions. The N-terminal domains of the two polypeptides with most similarity to mammalian PFK-2s contain a series of tandem repeat ankyrin motifs. In other proteins such motifs are known to mediate protein-protein interactions. Phylogenetic analysis suggests that the four different PFK-2/FBPase-2 isoenzymes found in Trypanosoma and Leishmania evolved from a single ancestral bifunctional enzyme within the trypanosomatid lineage. A possible explanation for the evolution of multiple monofunctional enzymes and for the presence of the ankyrin-motif repeats in the PFK-2 isoenzymes is presented.

Amino Acid Sequence↗

Evolutionary history of Oryza sativa LTR retrotransposons: a preliminary survey of the rice genome sequences.

BACKGROUND: LTR Retrotransposons transpose through reverse transcription of an RNA intermediate and are ubiquitous components of all eukaryotic genomes thus far examined. Plant genomes, in particular, have been found to be comprised of a remarkably high number of LTR retrotransposons. There is a significant body of direct and indirect evidence that LTR retrotransposons have contributed to gene and genome evolution in plants. RESULTS: To explore the evolutionary history of long terminal repeat (LTR) retrotransposons and their impact on the genome of Oryza sativa, we have extended an earlier computer-based survey to include all identifiable full-length, fragmented and solo LTR elements in the rice genome database as of April 2002. A total of 1,219 retroelement sequences were identified, including 217 full-length elements, 822 fragmented elements, and 180 solo LTRs. In order to gain insight into the chromosomal distribution of LTR-retrotransposons in the rice genome, a detailed examination of LTR-retrotransposon sequences on Chromosome 10 was carried out. An average of 22.3 LTR-retrotransposons per Mb were detected in Chromosome 10. CONCLUSIONS: Gypsy-like elements were found to be >4 x more abundant than copia-like elements. Eleven of the thirty-eight investigated LTR-retrotransposon families displayed significant subfamily structure. We estimate that at least 46.5% of LTR-retrotransposons in the rice genome are older than the age of the species (< 680,000 years). LTR-retrotransposons present in the rice genome range in age from those just recently inserted up to nearly 10 million years old. Approximately 20% of LTR retrotransposon sequences lie within putative genes. The distribution of elements across chromosome 10 is non-random with the highest density (48 elements per Mb) being present in the pericentric region.

Chromosomes, Plant↗

Nomenclature of the human immunoglobulin lambda (IGL) genes.

'Nomenclature of the Human Immunoglobulin Lambda (IGL) Genes', the 18th report of the 'IMGT Locus in Focus' section, provides the first complete list of all the human IGL genes. The total number of human IGL genes per haploid genome is 87--96 (93--102 if the orphons are included), of which 37--43 genes are functional. IMGT/Human Genome Organization (HUGO) gene names and definitions of the human IGL genes on chromosome 22q11.2 and IGL orphons on chromosomes 8 and 22 are provided with the gene functionality and the number of alleles, according to the rules of the IMGT Scientific chart, with the accession numbers of the IMGT reference sequences and with the accession ID of the Genome Database GDB and NCBI LocusLink databases, in which all the IMGT human IGL genes have been entered. The tables are available at the IMGT Marie-Paule page of IMGT, the international ImMunoGeneTics database (http://imgt.cines.fr) created by Marie-Paule Lefranc, Université Montpellier II, CNRS, France.

Alleles↗

Two thimet oligopeptidase-like Pz peptidases produced by a collagen-degrading thermophile, Geobacillus collagenovorans MO-1.

A collagen-degrading thermophile, Geobacillus collagenovorans MO-1, was found to produce two metallopeptidases that hydrolyze the synthetic substrate 4-phenylazobenzyloxycarbonyl-Pro-Leu-Gly-Pro-D-Arg (Pz-PLGPR), containing the collagen-specific sequence -Gly-Pro-X-. The peptidases, named Pz peptidases A and B, were purified to homogeneity and confirmed to hydrolyze collagen-derived oligopeptides but not collagen itself, indicating that Pz peptidases A and B contribute to collagen degradation in collaboration with a collagenolytic protease in G. collagenovorans MO-1. There were many similarities between Pz peptidases A and B in their catalytic properties; however, they had different molecular masses and shared no antigenic groups against the respective antibodies. Their primary structures clarified from the cloned genes showed lower identity (22%). From homology analysis for proteolytic enzymes in the database, the two Pz peptidases belong to the M3B family. In addition, Pz peptidases A and B shared high identities of over 70% with unassigned peptidases and oligopeptidase F-like peptidases of the M3B family, respectively. Those homologue proteins are putative in the genome database but form two distinct segments, including Pz peptidases A and B, in the phylogenic tree. Mammalian thimet oligopeptidases, which were previously thought to participate in collagen degradation and share catalytic identities with Pz peptidases, were found to have lower identities in the overall primary sequence with Pz peptidases A and B but a significant resemblance in the vicinity of the catalytic site.

Amino Acid Sequence↗

Phylogeny of the sex-determining gene Sex-lethal in insects.

The Sex-lethal (SXL) protein belongs to the family of RNA-binding proteins and is involved in the regulation of pre-mRNA splicing. SXL has undergone an obvious change of function during the evolution of the insect clade. The gene has acquired a pivotal role in the sex-determining pathway of Drosophila, although it does not act as a sex determiner in non-drosophilids. We collected SXL sequences of insect species ranging from the pea aphid (Acyrtho siphom pisum) to Drosophila melanogaster by searching published articles, sequencing cDNAs, and exploiting homology searches in public EST and whole-genome databases. The SXL protein has moderately conserved N- and C-terminal regions and a well-conserved central region including 2 RNA recognition motifs. Our phylogenetic analysis shows that a single orthologue of the Drosophila Sex-lethal (Sxl) gene is present in the genomes of the malaria mosquito Anopheles gambiae, the honeybee Apis mellifera, the silkworm Bombyx mori, and the red flour beetle Tribolium castaneum. The D. melanogaster, D. erecta, and D. pseudoobscura genomes, however, contain 2 paralogous genes, Sxl and CG3056, which are orthologous to the Anopheles, Apis, Bombyx, and Tribolium Sxl. Hence, a duplication in the fly clade generated Sxl and CG3056. Our hypothesis maintains that one of the genes, Sxl, adopted the new function of sex determiner in Drosophila, whereas the other, CG3056, continued to serve some or all of the yet-unknown ancestral functions.

Amino Acid Sequence↗

Multiple new and isolated families within the mouse superfamily of V1r vomeronasal receptors.

Seven-transmembrane-domain proteins encoded by the vomeronasal receptor V1r and V2r gene superfamilies, and expressed by vomeronasal sensory neurons, are believed to be pheromone receptors in rodents. Four V1r gene families have been described in the mouse (V1ra, V1rb, V1rc and V3r). Here we have screened near-complete mouse genomic databases to obtain a first global draft of the mouse V1r repertoire, including 104 new V1r genes. It comprises eight new and extremely isolated families in addition to the four families previously identified. Members of these new families were expressed in vomeronasal sensory neurons. The genome-wide view revealed great sequence diversity within the V1r superfamily. Phylogenetic analyses suggested an ancient original radiation, followed by the isolation, divergence and expansion of families by extensive gene duplications and frequent gene loss. The isolated nature of these gene families probably reflects a specialization of different receptor classes in the detection of specific types of chemicals.

Animals↗

Developmental roles of pufferfish Hox clusters and genome evolution in ray-fin fish.

The pufferfish skeleton lacks ribs and pelvic fins, and has fused bones in the cranium and jaw. It has been hypothesized that this secondarily simplified pufferfish morphology is due to reduced complexity of the pufferfish Hox complexes. To test this hypothesis, we determined the genomic structure of Hox clusters in the Southern pufferfish Spheroides nephelus and interrogated genomic databases for the Japanese pufferfish Takifugu rubripes (fugu). Both species have at least seven Hox clusters, including two copies of Hoxb and Hoxd clusters, a single Hoxc cluster, and at least two Hoxa clusters, with a portion of a third Hoxa cluster in fugu. Results support genome duplication before divergence of zebrafish and pufferfish lineages, followed by loss of a Hoxc cluster in the pufferfish lineage and loss of a Hoxd cluster in the zebrafish lineage. Comparative analysis shows that duplicate genes continued to be lost for hundreds of millions of years, contrary to predictions for the permanent preservation of gene duplicates. Gene expression analysis in fugu embryos by in situ hybridization revealed evolutionary change in gene expression as predicted by the duplication-degeneration-complementation model. These experiments rule out the hypothesis that the simplified pufferfish body plan is due to reduction in Hox cluster complexity, and support the notion that genome duplication contributed to the radiation of teleosts into half of all vertebrate species by increasing developmental diversification of duplicate genes in daughter lineages.

Animals↗

Novel sigmaF-dependent genes of Escherichia coli found using a specified promoter consensus.

Availability of whole genome information opens new bioinformatics approaches to study global regulation. We developed a program, named ScanProm, that allows to search a genome database for promoter consensus elements. The program uses a multiple alignment of previously identified components of a regulon as an input and generates a consensus profile. The profile is then optimized by adjusting the cutoff value for position-specific similarity assessment and used for a genome scan to search for unknown members of the regulon. The candidates obtained are scored by their similarity to the consensus profile. The ScanProm program was applied to search for novel members of the class III flagellar regulon of Escherichia coli. The search template included the previously defined 4 bp (-35) and 8 bp (-10) promoter elements, presumably recognized by the flagellar-specific sigmaF, with additional 4 bp at the 3' of the -35 consensus. The majority of highly scoring candidates obtained from the whole genome sequence scan were known class III genes, although several new genes were also identified. We tested 10 novel highly scoring candidate class III genes by cloning their promoter fragments into a fusion vector designed to monitor the transcriptional activity with lacZ. Two of these genes, b2737(ygbK) and ppdAB, were found to be dependent on FlhDC, the master regulator of the flagellar genes. The regulation of these genes by sigmaF was further confirmed by comparing their expression in the wild-type and fliA backgrounds. An overproduction or inactivation of these genes did not exhibit any notable phenotypes in motility or chemotaxis.

Base Sequence↗

Identification and characterization of a novel bacterial virulence factor that shares homology with mammalian Toll/interleukin-1 receptor family proteins.

Many important bacterial virulence factors act as mimics of mammalian proteins to subvert normal host cell processes. To identify bacterial protein mimics of components of the innate immune signaling pathway, we searched the bacterial genome database for proteins with homology to the Toll/interleukin-1 receptor (TIR) domain of the mammalian Toll-like receptors (TLRs) and their adaptor proteins. A previously uncharacterized gene, which we have named tlpA (for TIR-like protein A), was identified in the Salmonella enterica serovar Enteritidis genome that is predicted to encode a protein resembling mammalian TIR domains, We show that overexpression of TlpA in mammalian cells suppresses the ability of mammalian TIR-containing proteins TLR4, IL-1 receptor, and MyD88 to induce the transactivation and DNA-binding activities of NF-kappaB, a downstream target of the TIR signaling pathway. In addition, TlpA mimics the previously characterized Salmonella virulence factor SipB in its ability to induce activation of caspase-1 in a mammalian cell transfection model. Disruption of the chromosomal tlpA gene rendered a virulent serovar Enteritidis strain defective in intracellular survival and IL-1beta secretion in a cell culture infection model using human THP1 macrophages. Bacteria with disrupted tlpA also displayed reduced lethality in mice, further confirming an important role for this factor in pathogenesis. Taken together, our findings demonstrate that the bacterial TIR-like protein TlpA is a novel prokaryotic modulator of NF-kappaB activity and IL-1beta secretion that contributes to serovar Enteritidis virulence.

Amino Acid Sequence↗

Gene discovery by e-genetics: Drosophila odor and taste receptors.

A new algorithm that examines DNA databases for proteins that have a particular structure, as opposed to a particular sequence, represents a novel 'e-genetics' approach to gene discovery. The algorithm has successfully identified new G-protein-coupled receptors, which have a characteristic seven-transmembrane-domain structure, from the Drosophila genome database. In particular, it has revealed novel families of odor receptors and taste receptors, which had long eluded identification by other means. The two new gene families, the Or and Gr genes, are expressed in neurons of olfactory and taste sensilla and are highly divergent from all other known G-protein-coupled receptor genes. Modification of the algorithm should allow identification of other classes of multitransmembrane-domain protein.

Algorithms↗

Proteomic analysis of the human colon carcinoma cell line (LIM 1215): development of a membrane protein database.

The proteomic definition of plasma membrane proteins is an important initial step in searching for novel tumor marker proteins expressed during the different stages of cancer progression. However, due to the charge heterogeneity and poor solubility of membrane-associated proteins this subsection of the cell's proteome is often refractory to two-dimensional electrophoresis (2-DE), the current paradigm technology for studying protein expression profiles. Here, we describe a non-2-DE method for identifying membrane proteins. Proteins from an enriched membrane preparation of the human colorectal carcinoma cell line LIM1215 were initially fractionated by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE, 4-20%). The unstained gel was cut into 16 x 3 mm slices, and peptide mixtures resulting from in-gel tryptic digestion of each slice were individually subjected to capillary-column reversed phase-high performance liquid chromatography (RP-HPLC) coupled with electrospray ionization-ion trap-mass spectrometry (ESI-IT-MS). Interrogation of genomic databases with the resulting collision-induced dissociation (CID) generated peptide ion fragment data was used to identify the proteins in each gel slice. Over 284 proteins (including 92 membrane proteins) were identified, including many integral membrane proteins not previously identified by 2-DE, many proteins seen at the genomic level only, as well as several proteins identified by expressed sequence tags (ESTs) only. Additionally, a number of peptides, identified by de novo MS sequence analysis, have not been described in the databases. Further, a "targeted" ion approach was used to unambiguously identify known low-abundance plasma membrane proteins, using the membrane-associated A33 antigen, a gastrointestinal-specific epithelial cell protein, as an example. Following localization of the A33 antigen in the gel by immunoblotting, ions corresponding to the theoretical A33 antigen tryptic peptide masses were selected using an "inclusion" mass list for automated sequence analysis. Six peptides corresponding to the A33 antigen, present at levels well below those accessible using the standard automated "nontargeted" approach, were identified. The membrane protein database may be accessed via the World Wide Web (WWW) at http://www.ludwig. edu.au/jpsl/jpslhome.html.

Colonic Neoplasms↗

cDNA cloning of myosin heavy chain genes from medaka Oryzias latipes embryos and larvae and their expression patterns during development.

Several sarcomeric myosin heavy chains (MYHs) were cloned from embryos and larvae of medaka Oryzias latipes. Three genes encoding medaka MYHs (mMYHs) predominantly expressed in embryos (mMYH(emb1)) and larvae (mMYH(L1) and mMYH(L2)), all belonged to fast skeletal MYHs, showing spatiotemporally different expression patterns during development. Besides these mMYHs, a few novel mMYHs were cloned from embryos and larvae at hatching. Whereas mMYH(emb2), mMYH(emb3), and mMYH(L3) belonged to fast skeletal MYH, mMYH(C1) and mMYH(C2) did to slow/cardiac MYH. mMYH(emb1) was expressed ahead of mMYH(L1) and mMYH(L2). In situ hybridization analysis demonstrated that the transcripts of mMYH(emb1) and mMYH(C1) were located in the horizontal myoseptum, whereas those of mMYH(L1) and mMYH(L2) in the inner part of myotomes and pharyngeal muscles, and those of mMYH(C2) in the heart rudiment. In silico cloning based on the medaka genome database showed another mMYHs of the slow/cardiac types, mMYH(C3) and mMYH(C4).

Animals↗

Molecular cloning of rat Spetex2 family genes mapped on chromosome 15p16, encoding a 23-kilodalton protein associated with the plasma membranes of haploid spermatids.

We used differential display in combination with cDNA cloning to isolate a novel rat gene, designated as Spetex2, that has an open reading frame of 582 nucleotides, encoding a protein of 194 amino acids. Spetex2 mRNA was highly expressed in testis and spleen, and its expression in rat testis was developmentally up-regulated. In situ hybridization revealed that Spetex2 mRNA was predominantly expressed in haploid spermatids at steps 1-13 within the seminiferous epithelium. A BLAST search against rat genome databases at the National Center for Biotechnology Information revealed that the Spetex2 gene is composed of four exons and is mapped to at least 18 loci in a cluster on rat chromosome 15p16, indicating that the genes occur as a repeated tandem array over a long stretch of genomic DNA. By immunocytochemical analysis with confocal laser-scanning microscopy, SPETEX2 protein was detected as a dot-like distribution on the cell periphery of haploid spermatids (steps 1-13) but was not observed in other spermatogenic cells. On the basis of these data, we hypothesize that SPETEX2 might be correlated with cell differentiation of spermaytids in rat testis.

Animals↗