Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Identification of glycosylphosphatidylinositol-anchored proteins in Arabidopsis. A proteomic and genomic analysis.

In a recent bioinformatic analysis, we predicted the presence of multiple families of cell surface glycosylphosphatidylinositol (GPI)-anchored proteins (GAPs) in Arabidopsis (G.H.H. Borner, D.J. Sherrier, T.J. Stevens, I.T. Arkin, P. Dupree [2002] Plant Physiol 129: 486-499). A number of publications have since demonstrated the importance of predicted GAPs in diverse physiological processes including root development, cell wall integrity, and adhesion. However, direct experimental evidence for their GPI anchoring is mostly lacking. Here, we present the first, to our knowledge, large-scale proteomic identification of plant GAPs. Triton X-114 phase partitioning and sensitivity to phosphatidylinositol-specific phospholipase C were used to prepare GAP-rich fractions from Arabidopsis callus cells. Two-dimensional fluorescence difference gel electrophoresis and one-dimensional sodium dodecyl sulfate-polyacrylamide gel electrophoresis demonstrated the existence of a large number of phospholipase C-sensitive Arabidopsis proteins. Using liquid chromatography-tandem mass spectrometry, 30 GAPs were identified, including six beta-1,3 glucanases, five phytocyanins, four fasciclin-like arabinogalactan proteins, four receptor-like proteins, two Hedgehog-interacting-like proteins, two putative glycerophosphodiesterases, a lipid transfer-like protein, a COBRA-like protein, SKU5, and SKS1. These results validate our previous bioinformatic analysis of the Arabidopsis protein database. Using the confirmed GAPs from the proteomic analysis to train the search algorithm, as well as improved genomic annotation, an updated in silico screen yielded 64 new candidates, raising the total to 248 predicted GAPs in Arabidopsis.

Arabidopsis↗

An update of the HLA genomic region, locus information and disease associations: 2004.

The human major histocompatibility (MHC) genomic region at chromosomal position 6p21 encodes the six classical transplantation HLA genes and many other genes that have important roles in the regulation of the immune system as well as in some fundamental cellular processes. This small segment of the human genome has been associated with more than 100 diseases, including common diseases--such as diabetes, rheumatoid arthritis, psoriasis, asthma and various autoimmune disorders. The MHC 3.6 Mb genomic sequence was first reported in 1999 with the annotation of 224 gene loci. The locus and allelic information of the MHC continue to be updated by identifying newly mapped expressed genes and pseudogenes based on comparative genomics, SNP analysis and cDNA projects. Since 1999, new innovations in bioinformatics and gene-specific functional databases and studies on the MHC genes have resulted in numerous changes to gene names and better ways to update and link the MHC gene symbols, names and sequences together with function, variation and disease associations. In this study, we present a brief overview of the MHC genomic structure and the recent information that we have gathered on the MHC gene loci via LocusLink at the National Centre for Biological Information (http://www.ncbi.nih.gov/.) and the MHC genes' association with various diseases taken from publications and records in public databases, such as the Online Mendelian Inheritance in Man and the Genetic Association Database.

Computational Biology↗

Bioinformatic identification of tandem repeat antigens of the Leishmania donovani complex.

With large amounts of parasite gene sequence available, additional bioinformatic tools to screen these sequences for identifying genes encoding antigens are needed. Proteins containing tandem repeat (TR) domains are often B-cell antigens, and antibody responses toward TR domains of the proteins are dominant in human infected with certain parasites. We hypothesized that antigens of serological significance could be identified with a search for TR domains. Here we show the result of bioinformatic screening of the gene sequence database of the parasitic protozoan Leishmania infantum. Of 8,191 genes scanned, 64 genes contained TR domains. Of the 64 genes, 22 encoded previously characterized antigens; the remaining 42 genes were previously uncharacterized. By using sera from Sudanese visceral leishmaniasis patients, we confirmed that the TR domains of LinJ11.0070, LinJ25.1100, LinJ27.0400, and LinJ29.0110, which were from the 42 uncharacterized proteins, are also antigenic. The results suggest the validity of this approach for identifying leishmanial antigens of serological significance.

Animals↗

The cobZ gene of Methanosarcina mazei Go1 encodes the nonorthologous replacement of the alpha-ribazole-5'-phosphate phosphatase (CobC) enzyme of Salmonella enterica.

Open reading frame (ORF) Mm2058 of the methanogenic archaeon Methanosarcina mazei strain Gö1 was shown in vivo and in vitro to encode the nonorthologous replacement of the alpha-ribazole-phosphate phosphatase (CobC; EC 3.1.3.73) enzyme of Salmonella enterica serovar Typhimurium LT2. Bioinformatics analysis of sequences available in databases tentatively identified ORF Mm2058, which was cloned under the control of an inducible promoter and was used to support growth of an S. enterica strain under conditions that demanded CobC-like activity. The Mm2058 protein was expressed with a decahistidine tag at its N terminus and was purified to homogeneity using nickel affinity chromatography. High-performance liquid chromatography followed by electrospray ionization mass spectrometry showed that the Mm2058 protein had phosphatase activity that converted alpha-ribazole-5'-phosphate to alpha-ribazole, as reported for the bacterial CobC enzyme. On the basis of the data reported here, we refer to ORF Mm2058 as cobZ. We tested the prediction by Rodionov et al. (D. A. Rodionov, A. G. Vitreschak, A. A. Mironov, and M. S. Gelfand, J. Biol. Chem. 278:41148-41159, 2003) that ORF HSL01294 (also called Vng1577) encoded the nonorthologous replacement of the bacterial CobC enzyme in the extremely halophilic archaeon Halobacterium sp. strain NRC-1. A strain of the latter carrying an in-frame deletion of ORF Vng1577 was not a cobalamin auxotroph, suggesting that either there is redundancy of this function in Halobacterium or the gene was misannotated.

Bacterial Proteins↗

Serial analysis of gene expression (SAGE) in bovine trypanotolerance: preliminary results.

In Africa, trypanosomosis is a tsetse-transmitted disease which represents the most important constraint to livestock production. Several indigenous West African taurine Bos taurus) breeds, such as the Longhorn (N'Dama) cattle are well known to control trypanosome infections. This genetic ability named "trypanotolerance" results from various biological mechanisms under multigenic control. The methodologies used so far have not succeeded in identifying the complete pool of genes involved in trypanotolerance. New post genomic biotechnologies such as transcriptome analyses are efficient in characterising the pool of genes involved in the expression of specific biological functions. We used the serial analysis of gene expression (SAGE) technique to construct, from Peripheral Blood Mononuclear Cells of an N'Dama cow, 2 total mRNA transcript libraries, at day 0 of a Trypanosoma congolense experimental infection and at day 10 post-infection, corresponding to the peak of parasitaemia. Bioinformatic comparisons in the bovine genomic databases allowed the identification of 187 up- and down- regulated genes, EST and unknown functional genes. Identification of the genes involved in trypanotolerance will allow to set up specific microarray sets for further metabolic and pharmacological studies and to design field marker-assisted selection by introgression programmes.

Africa South of the Sahara↗

Use of the Serial Analysis of Gene Expression (SAGE) method in veterinary research: A concrete application in the study of the bovine trypanotolerance genetic control.

New postgenomic biotechnologies, such as transcriptome analyses, are now able to characterize the full complement of genes involved in the expression of specific biological functions. One of these is the Serial Analysis of Gene Expression (SAGE) technique, which consists of the construction of transcripts libraries for a quantitative analysis of the entire gene(s) expressed or inactivated at a particular step of cellular activation. Bioinformatic comparisons in the bovine genomic databases allow the identification of several up- and downregulated genes, expressed sequence tags, and unknown functional genes directly involved in the genetic control of the studied biological mechanism. We present and discuss the preliminary results in comparing the expressed genes in two total mRNA transcripts libraries obtained during an experimental Trypanosoma congolense infection in one trypanotolerant N'Dama animal cow. Knowing all the functional genes involved in the trypanotolerance control will permit validation of some results obtained with the quantitative trait locus approach, to set up specific microarrays sets for further metabolic and pharmacological studies, and to design field marker-assisted selection by introgression programs.

Animals↗

Overexpression of TCF7L2 promotes the viability and migration of MHCC-97H human hepatocellular carcinoma cells by upregulating MT-ND4L.

BACKGROUND: Hepatocellular carcinoma (HCC) is a highly aggressive cancer with high metabolic adaptability. TCF7L2, a transcription factor implicated in type 2 diabetes and cancer, is overexpressed in HCC. However, its specific role in HCC metabolic reprogramming is not well defined. We aimed to elucidate the previously unrecognized molecular mechanisms through which TCF7L2 impacts HCC progression. METHODS: To investigate the function of TCF7L2, a stable MHCC-97H cell line with TCF7L2 overexpression was established via lentiviral transduction. Cell viability and migration were assessed by Cell Counting Kit-8 (CCK-8) and Transwell assays. Transcriptomic profiling [RNA sequencing (RNA-seq)] was performed to identify differentially expressed genes (DEGs). Functional enrichment analysis [Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), Gene Set Enrichment Analysis (GSEA)] and bioinformatics promoter analysis (the JASPAR CORE database) were conducted. Clinical correlations, survival analysis, and tumor microenvironment (TME) interrogation were performed using The Cancer Genome Atlas Liver Hepatocellular Carcinoma (TCGA-LIHC) cohort and single-cell datasets [Human Protein Atlas (HPA), CellChat]. Drug sensitivity was predicted via the Genomics of Drug Sensitivity in Cancer (GDSC) database. RESULTS: TCF7L2 overexpression significantly promoted HCC cell proliferation and migration. Transcriptomic analysis revealed that TCF7L2 drives a profound metabolic shift, with key enrichments in lipid homeostasis, fatty acid β-oxidation, and the PI3K/Akt pathway. Mechanistically, TCF7L2 directly binds to the promoter of CPT1A, the rate-limiting enzyme of fatty acid oxidation, and indirectly upregulates the mitochondrial gene MT-ND4Lvia a strong positive correlation with the mitochondrial transcription factor TFAM. In clinical cohorts, TCF7L2 was overexpressed in HCC and its expression correlated positively with MT-ND4L, MKI67, and SNAI1, and served as a predictor of poor overall survival (OS). Furthermore, TCF7L2-high tumors were enriched in hepatic progenitor cell (HPC)-like niches, mediated by enhanced ANGPTL4 signaling. High TCF7L2 expression predicted increased sensitivity to PI3K/mTOR pathway inhibitors. CONCLUSIONS: TCF7L2 acts as a master metabolic regulator in HCC, coordinating lipid catabolism and mitochondrial biogenesis to drive aggressive tumor behavior. It further remodels the TME towards an HPC-like state and predicts sensitivity to metabolic-targeted therapies. These findings identify TCF7L2 as a key prognostic biomarker and a promising therapeutic target.

MHCC-97H hepatocellular carcinoma cells (MHCC-97H ↗

On the way to understand biological complexity in plants: S-nutrition as a case study for systems biology.

The establishment of technologies for high-throughput DNA sequencing (genomics), gene expression (transcriptomics), metabolite and ion analysis (metabolomics/ionomics) and protein analysis (proteomics) carries with it the challenge of processing and interpreting the accumulating data sets. Publicly accessible databases and newly development and adapted bioinformatic tools are employed to mine this data in order to filter relevant correlations and create models describing physiological states. These data allow the reconstruction of networks of interactions of the various cellular components as enzyme activities and complexes, gene expression, metabolite pools or pathway flux modes. Especially when merging information from transcriptomics, metabolomics and proteomics into consistent models, it will be possible to describe and predict the behaviour of biological systems, for example with respect to endogenous or environmental changes. However, to capture the interactions of network elements requires measurements under a variety of conditions to generate or refine existing models. The ultimate goal of systems biology is to understand the molecular principles governing plant responses and consistently explain plant physiology.

Arabidopsis↗

Differentiation of normal skin and melanoma using high resolution hyperspectral imaging.

We investigated the use of high resolution hyperspectral imaging microscopy to detect abnormalities in skin tissue using hematoxylin eosin stained preparations of normal and abnormal skin, benign nevi and melanomas. A goal of this study was to provide objective data that could be utilized by any researcher; and form the beginnings of a reference spectral data base. All spectral characterizations were acquired in percent transmission, and absorption, with contiguous wavelength acquisition between 400 and 800 nm; and a spectral resolution of approximately 1 nm. Biopsy sections were characterized with varying sample thickness, staining and magnification in order to determine their impact on spectral characterizations. Spectra were classified using spectral waveform cross correlation analysis, an algorithm that is linearity invariant. Classified spectra were incorporated into spectral libraries; and all spectra acquired from the field of view were correlated with library spectra to a quantified, user determined, confidence threshold (minimum correlation coefficient). The results revealed that all skin conditions in our initial data sets could be objectively differentiated providing that staining and section thickness was controlled. We also demonstrated that it is likely that a reference spectral library database could be created to include bioinformatics and cluster analysis. This would assist multiple laboratories to participate in the input and retrieval of target spectral information.

Biopsy↗

Proteomic technology and its biomedical applications.

Proteomics has its origins in two-dimensional gel electrophoresis (2-DE), a technique developed more than twenty years ago. 2-DE has a high-resolution capacity, and was initially used primarily for separating and characterizing proteins in complex mixtures. 2-DE remains an important tool for protein identification, but is now normally coupled with mass spectrometry (MS), a technique which has advanced considerably in recent years. The recent completion of human genome project has produced a large DNA database which can be utilized through bioinformatics, and the next challenge for scientists is to uncover the entire proteome of a particular organism. The integration of genomic and proteomic data will help to elucidate the functions of proteins in the pathogenesis of diseases and the ageing process, and could lead to the discovery of novel drug target proteins and biomarkers of diseases. This review describes recent advances in proteomic technology and discusses the potential applications of proteomics in biomedical research.

Biomedical Research↗

The Type 1 Diabetes Genetics Consortium.

The Type 1 Diabetes Genetics Consortium (T1DGC) is an international, multicenter program organized to promote research to identify genes and their alleles that determine an individual's risk for type 1 diabetes (T1D). The primary goal of the T1DGC is to establish resources and data that can be used by, and that is fully accessible to, the research community in the study of T1D. All the information on T1DGC can be accessed at the following web address: http://www.t1dgc.org. A resource base of well-characterized families is being assembled that will facilitate the localization and characterization of T1D susceptibility genes. From these families, the T1DGC is establishing banks of DNA, serum, plasma, and cell lines, as well as useful databases. The T1DGC also sponsors training opportunities (bioinformatics) and technology transfer (HLA genotyping).

Alleles↗

Bioinformatic and experimental tools for identification of single-nucleotide polymorphisms in genes with a potential role for the development of the insulin resistance syndrome.

OBJECTIVES: Genes with a possible role for the development of the insulin resistance syndrome (IRS) were scanned for novel single-nucleotide polymorphisms (SNPs) using bioinformatics. METHODS: GenBank mRNA sequences were compared to the human EST database using gapped BLAST, software that is available on the internet. Mismatches between the search and the EST sequences indicated potential SNPs. Thirty-two SNPs in 13 genes were randomly chosen for experimental verification. PCR and direct sequencing were used to determine the 'true' SNPs. A random sample of 30 Swedish men with slightly elevated diastolic blood pressure (85-94 mmHg) obtained from a population-based study was selected for the sequencing. After completion of these stages, the potential SNPs were checked against the large and rapidly expanding SNP databases HGBASE and NCBI. RESULTS: EST searches of 146 genes revealed 106 potential SNPs in 44 genes. Experimental analysis of 32 of these potential SNPs verified two SNPs; endothelin receptor A 1471 G/C (3' UTR) and PAI-1 Trp514Arg from a T/C exchange. These two SNPs were also identified in the NCBI and HGBASE databases together with two polymorphisms that were not experimentally identified in our homogeneous Swedish population. Overall, the HGBASE and NCBI databases contained entries of 22% (23 out of 106) of the SNPs identified through our EST searches. CONCLUSIONS: In the search for genetic variations causing complex diseases like IRS in homogeneous populations (such as the Swedish one used here), important information can be obtained through bioinformatic searches of human genome databases and experimental verification.

Adult↗

FastGroupII: a web-based bioinformatics platform for analyses of large 16S rDNA libraries.

BACKGROUND: High-throughput sequencing makes it possible to rapidly obtain thousands of 16S rDNA sequences from environmental samples. Bioinformatic tools for the analyses of large 16S rDNA sequence databases are needed to comprehensively describe and compare these datasets. RESULTS: FastGroupII is a web-based bioinformatics platform to dereplicate large 16S rDNA libraries. FastGroupII provides users with the option of four different dereplication methods, performs rarefaction analysis, and automatically calculates the Shannon-Wiener Index and Chao1. FastGroupII was tested on a set of 16S rDNA sequences from coral-associated Bacteria. The different grouping algorithms produced similar, but not identical, results. This suggests that 16S rDNA datasets need to be analyzed in multiple ways when being used for community ecology studies. CONCLUSION: FastGroupII is an effective bioinformatics tool for the trimming and dereplication of 16S rDNA sequences. Several standard diversity indices are calculated, and the raw sequences are prepared for downstream analyses.

Algorithms↗

MX1 promotes gastric cancer cell migration via inhibiting ANXA2 ubiquitination and degradation.

Gastric cancer (GC) is a globally lethal malignancy, with invasion and metastasis driving treatment failure and poor prognosis. MX dynamin like GTPase 1 (MX1) shows tumor-specific functional heterogeneity, while its expression, biological functions and molecular mechanisms in GC remain unclear. Here, we explored MX1's clinical significance and its regulatory mechanism in GC cell migration. We integrated public databases and institutional paired clinical samples for bioinformatics analysis of MX1's correlation with clinical outcomes, and verified its pro-migratory effect via Transwell and wound healing assays. Co-immunoprecipitation/mass spectrometry (Co-IP/MS), immunofluorescence and ubiquitination assays were used to identify MX1-interacting proteins and dissect the underlying mechanism, and the Genomics of Drug Sensitivity in Cancer database was applied for chemosensitivity analysis. MX1 was aberrantly upregulated in GC tissues and served as an independent prognostic biomarker, with high expression associated with shortened overall, first-progression and post-progression survival. MX1 promoted GC cell migration and epithelial-mesenchymal transition pathway enrichment, and directly bound Annexin A2 (ANXA2) in the cytoplasm; both were co-enriched in endothelial and epithelial cells by single-cell sequencing. MX1 dose-dependently upregulated ANXA2 protein (without affecting its mRNA) by inhibiting NEDD4L/TRIM65-mediated ANXA2 ubiquitination and degradation, enhancing ANXA2 stability. Additionally, high MX1 expression correlated with increased paclitaxel sensitivity in GC patients based on database analysis, and CCK-8 assays confirmed that MX1 overexpression significantly reduced the paclitaxel IC50 in gastric cancer cells, supporting its potential as a predictive biomarker for paclitaxel efficacy. This study demonstrates that MX1 promotes GC cell migration by suppressing ANXA2 ubiquitination and degradation, highlighting the critical role of the MX1-ANXA2 axis in GC progression. These findings provide novel molecular targets and theoretical support for GC prognostic evaluation, individualized chemotherapy and targeted therapy.

ANXA2↗

Human protein reference database as a discovery resource for proteomics.

The rapid pace at which genomic and proteomic data is being generated necessitates the development of tools and resources for managing data that allow integration of information from disparate sources. The Human Protein Reference Database (http://www.hprd.org) is a web-based resource based on open source technologies for protein information about several aspects of human proteins including protein-protein interactions, post-translational modifications, enzyme-substrate relationships and disease associations. This information was derived manually by a critical reading of the published literature by expert biologists and through bioinformatics analyses of the protein sequence. This database will assist in biomedical discoveries by serving as a resource of genomic and proteomic information and providing an integrated view of sequence, structure, function and protein networks in health and disease.

Computational Biology↗

SNPeffect v2.0: a new step in investigating the molecular phenotypic effects of human non-synonymous SNPs.

UNLABELLED: Single nucleotide polymorphisms (SNPs) constitute the most fundamental type of genetic variation in human populations. About 75 000 of these reported variations cause an amino acid change in the translated protein. An important goal in genomic research is to understand how this variability affects protein function, and whether or not particular SNPs are associated to disease susceptibility. Accordingly, the SNPeffect database uses sequence- and structure-based bioinformatics tools to predict the effect of non-synonymous SNPs on the molecular phenotype of proteins. SNPeffect analyses the effect of SNPs on three categories of functional properties: (1) structural and thermodynamic properties affecting protein dynamics and stability (2) the integrity of functional and binding sites and (3) changes in posttranslational processing and cellular localization of proteins. The search interface of the database can be used to search specifically for polymorphisms that are predicted to cause a change in one of these properties. Now based on the Ensembl human databases, the SNPeffect database has been remodeled to better fit an automatically updatable structure. The current edition holds the molecular phenotype of 74 567 nsSNPs in 23 426 proteins. AVAILABILITY: SNPeffect can be accessed through http://snpeffect.vib.be.

Algorithms↗