Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

A glycoproteome database of normal human liver tissue.

PURPOSE: To extensively investigate the glycoproteins of normal human liver tissue, constructing the glycoprotein profile and database of the normal human liver tissue. METHODS: The total proteins were extracted from the normal human liver tissue and then subjected to two-dimensional electrophoresis (2-DE). Finally, 2-DE gels were stained according to the methods of multiplexed proteomics (MP) technology. Glycoprotein spots were excised from 2-DE gel and then characterized by matrix assisted laser desorption/ionization-time of flight mass spectrometry (MALDI-TOF-MS). RESULTS: The PDQuest software detected 1,011 glycoprotein spots and 1,923 total protein spots in the 2-DE gels of sample from the normal human liver tissue. Furthermore, 116 species of glycoproteins were successfully identified via peptide mass profiling using MALDI-TOF-MS/MS and annotated to our databases. In addition, we also applied bioinformatics softwares to predict N- or O-glycosylation sites of identified glycoproteins. CONCLUSION: This study demonstrates the feasibility of a novel technological platform to contruct glycoprotein databases. These results lay the foundation for future physiological and pathological studies of the human liver.

Databases, Protein↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database is a comprehensive database of DNA and RNA sequences directly submitted from researchers and genome sequencing groups and collected from the scientific literature and patent applications. In collaboration with DDBJ and GenBank the database is produced, maintained and distributed at the European Bioinformatics Institute (EBI) and constitutes Europe's primary nucleotide sequence resource. Database releases are produced quarterly and are distributed on CD-ROM. EBI's network services allow access to the most up-to-date data collection via Internet and World Wide Web interface, providing database searching and sequence similarity facilities plus access to a large number of additional databases.

Academies and Institutes↗

Genomic scale sub-family assignment of protein domains.

Many classification schemes for proteins and domains are either hierarchical or semi-hierarchical yet most databases, especially those offering genome-wide analysis, only provide assignments to sequences at one level of their hierarchy. Given an established hierarchy, the problem of assigning new sequences to lower levels of that existing hierarchy is less hard (but no less important) than the initial top level assignment which requires the detection of the most distant relationships. A solution to this problem is described here in the form of a new procedure which can be thought of as a hybrid between pairwise and profile methods. The hybrid method is a general procedure that can be applied to any pre-defined hierarchy, at any level, including in principle multiple sub-levels. It has been tested on the SCOP classification via the SUPERFAMILY database and performs significantly better than either pairwise or profile methods alone. Perhaps the greatest advantage of the hybrid method over other possible approaches to the problem is that within the framework of an existing profile library, the assignments are fully automatic and come at almost no additional computational cost. Hence it has already been applied at the SCOP family level to all genomes in the SUPERFAMILY database, providing a wealth of new data to the biological and bioinformatics communities.

Computational Biology↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl/), maintained at the European Bioinformatics Institute (EBI), incorporates, organizes and distributes nucleotide sequences from public sources. The database is a part of an international collaboration with DDBJ (Japan) and GenBank (USA). Data are exchanged between the collaborating databases on a daily basis to achieve optimal synchrony. The web-based tool, Webin, is the preferred system for individual submission of nucleotide sequences, including Third Party Annotation (TPA) and alignment data. Automatic submission procedures are used for submission of data from large-scale genome sequencing centres and from the European Patent Office. Database releases are produced quarterly. The latest data collection can be accessed via FTP, email and WWW interfaces. The EBI's Sequence Retrieval System (SRS) integrates and links the main nucleotide and protein databases as well as many other specialist molecular biology databases. For sequence similarity searching, a variety of tools (e.g. FASTA and BLAST) are available that allow external users to compare their own sequences against the data in the EMBL Nucleotide Sequence Database, the complete genomic component subsection of the database, the WGS data sets and other databases. All available resources can be accessed via the EBI home page at http://www.ebi.ac.uk.

Animals↗

The physics and bioinformatics of binding and folding-an energy landscape perspective.

It has been recognized in the last few years that unstructured proteins play an important role in biological organisms, often participating in signal transduction, transcriptional regulation, and a variety of other regulatory activities. Various hypotheses have been put forward for the ubiquity of the unfolded state; rapid turnover, faster or more specific binding kinetics, multifunctionality may all possibly explain apparent ubiquitousness of unfolded proteins in eukaryotic cells. In this paper we extend the energy landscape theory of protein folding to construct an analytical model of how binding and folding are coupled thermodynamically when the energy landscape is partially rugged. To deduce the parameters that enter the theory, which is based on Generalized Random Energy Model, we have analyzed in a bioinformatic sense a large structural database of more than 500 protein complexes. We find that Miyazawa-Jernigan contact potential shows similar energy gaps for folding for both hydrophobic and hydrophilic proteins, but that for binding contacts hydrophobic interfaces turn out to be funneled while hydrophilic ones are antifunneled. This suggests evolution has found a mechanism for avoiding frustration between folding and binding by making use of indirect water-mediated interactions. By juxtaposing the monomeric protein folding free energy profile in the protein complex database with another database consisting of only well-folded monomers, we estimate that at least 15% of monomers in the former database are unfolded in the absence of partner protein interface interactions. When employing the parameters characteristic of these unfolded monomers to construct binding/folding phase diagrams, we find that these monomers would indeed fold if sufficiently stabilizing binding contacts, consistent with that fold, are formed.

Computational Biology↗

Glycosylation patterns of human chorionic gonadotropin revealed by liquid chromatography-mass spectrometry and bioinformatics.

Due to their extensive structural heterogeneity, the elucidation of glycosylation patterns in glycoproteins such as the subunits of human chorionic gonadotropin (hCG), hCG-alpha, and hCG-beta, remains one of the most challenging problems in the proteomic analysis of post-translational modifications. In consequence, glycosylation is usually studied after decomposition of the intact proteins to the proteolytic peptide level. However, by this approach all information about the combination of the different glycopeptides in the intact protein is lost. In this study we have, therefore, attempted to combine the results of glycan identification after tryptic digestion with molecular mass measurements on the native starting material of the new first WHO Reference Reagents (RR) for hCG-alpha (99/720) and hCG-beta (99/650). Despite the extremely high number of possible combinations of the glycans identified in the tryptic peptides by HPLC-MS (>1000 for hCG-alpha and >10 000 for hCG-beta), the mass spectra of intact hCG-alpha and hCG-beta revealed only a limited number of glycoforms present in hCG preparations from pools of pregnancy urines. Peak annotations for hCG-alpha were performed with the help of a bioinformatic algorithm that generated a database containing all possible modifications of the proteins, including modifications possibly introduced during sample preparation such as oxidation or truncation, for subsequent searches for combinations fitting the mass difference between the polypeptide backbone and the measured molecular masses. Fourteen different glycoforms of hCG-alpha, containing biantennary, partly sialylized hybrid-type glycans, including methionine-oxidized and N-terminally truncated forms, were identified. Mass spectra of high quality were also obtained for hCG-beta, however, a database search mass accuracy of +/-5 Da was insufficient to unambiguously assign the possible combinations of post-translational modifications. In summary, mass spectrometric fingerprints of intact molecules were shown to be highly useful for the characterization of glycosylation patterns of different hCG preparations such as the new first WHO RR for immunoassays and could be the first step in establishing biophysical reference methods for hCG and related molecules.

Chorionic Gonadotropin↗

Applications of quantitative digital image analysis to breast cancer research.

Our studies of radiogenic carcinogenesis in mouse and human models of breast cancer are based on the view that cell phenotype, microenvironment composition, communication between cells and within the microenvironment are important factors in the development of breast cancer. This is complicated in the mammary gland by its postnatal development, cyclic evolution via pregnancy and involution, and dynamic remodeling of epithelial-stromal interactions, all of which contribute to breast cancer susceptibility. Microscopy is the tool of choice to examine cells in context. Specific features can be defined using probes, antibodies, immunofluorescence, and image analysis to measure protein distribution, cell composition, and genomic instability in human and mouse models of breast cancer. We discuss the integration of image acquisition, analysis, and annotation to efficiently analyze large amounts of image data. In the future, cell and tissue image-based studies will be facilitated by a bioinformatics strategy that generates multidimensional databases of quantitative information derived from molecular, immunological, and morphological probes at multiple resolutions. This approach will facilitate the construction of an in vivo phenotype database necessary for understanding when, where, and how normal cells become cancer.

Animals↗

Metabolomics: building on a century of biochemistry to guide human health.

Medical diagnosis and treatment efficacy will improve significantly when a more personalized system for health assessment is implemented. This system will require diagnostics that provide sufficiently detailed information about the metabolic status of individuals such that assay results will be able to guide food, drug and lifestyle choices to maintain or improve distinct aspects of health without compromising others. Achieving this goal will use the new science of metabolomics - comprehensive metabolic profiling of individuals linked to the biological understanding of human integrative metabolism. Candidate technologies to accomplish this goal are largely available, yet they have not been brought into practice for this purpose. Metabolomic technologies must be sufficiently rapid, accurate and affordable to be routinely accessible to both healthy and acutely ill individuals. The use of metabolomic data to predict the health trajectories of individuals will require bioinformatic tools and quantitative reference databases. These databases containing metabolite profiles from the population must be built, stored and indexed according to metabolic and health status. Building and annotating these databases with the knowledge to predict how a specific metabolic pattern from an individual can be adjusted with diet, drugs and lifestyle to improve health represents a logical application of the biochemistry knowledge that the life sciences have produced over the past 100 years.

Journal Article↗

Biosynthesis of lysine in plants: evidence for a variant of the known bacterial pathways.

With the aim of elucidating how plants synthesize lysine, extracts prepared from corn, tobacco, Chlamydomonas and soybean were tested and found to lack detectable amounts of N-alpha-acyl-L,L-diaminopimelate deacylase or N-succinyl-alpha-amino-epsilon-ketopimelate-glutamate aminotransaminase, two key enzymes in the central part of the bacterial pathway for lysine biosynthesis. Corn extracts missing two key enzymes still carried out the overall synthesis of lysine when provided with dihydrodipicolinate. An analysis of available plant DNA sequences was performed to test the veracity of the negative biochemical findings. Orthologs of dihydrodipicolinate reductase and diaminopimelate epimerase (enzymes on each side of the central pathway) were readily found in the Arabidopsis thaliana genome. Orthologs of the known enzymes needed to convert tetrahydrodipicolinate to diaminopimelic acid (DAP) were not detected in Arabidopsis or in the plant DNA sequence databases. The biochemical and reinforcing bioinformatics results provide evidence that plants may use a novel variant of the bacterial pathways for lysine biosynthesis.

Bacteria↗

Chemokines in immunity.

Chemokines are a superfamily of small, heparin-binding cytokines that induce directed migration of various types of leukocytes through interactions with a group of seven-transmembrane G protein-coupled receptors. At present, over 40 members have been identified in humans. Until a few years ago, chemokines were mainly known as potent attractants for leukocytes such as neutrophils and monocytes, and were thus mostly regarded as the mediators of acute and chronic inflammatory responses. They had highly complex ligand-receptor relationships and their genes were regularly mapped on chromosomes 4 and 17 in humans. Recently, novel chemokines have been identified in rapid succession, mostly through application of bioinformatics on expressed sequence tag databases. A number of surprises have followed the identification of novel chemokines. They are constitutively expressed in lymphoid and other tissues with individually characteristic patterns. Most of them turned out to be highly specific for lymphocytes and dendritic cells. They have much simpler ligand-receptor relationships, and their genes are mapped to chromosomal loci different from the traditional chemokine gene clusters. Thus, the emerging chemokines are functionally and genetically quite different from the classical "inflammatory chemokines" and may be classified as "immune (system) chemokines" because of their profound importance in the genesis, homeostasis and function of the immune system. The emergence of immune chemokines has brought about a great deal of impact on the current immunological research, leading us to a better understanding on the fine traffic regulation of lymphocytes and dendritic cells. The immune chemokines and their receptors are also likely to be important future targets for therapeutic intervention of our immune responses.

Animals↗

Characterization of the omega class of glutathione transferases.

The Omega class of cytosolic glutathione transferases was initially recognized by bioinformatic analysis of human sequence databases, and orthologous sequences were subsequently discovered in mouse, rat, pig, Caenorhabditis elegans, Schistosoma mansoni, and Drosophila melanogaster. In humans and mice, two GSTO genes have been recognized and their genetic structures and expression patterns identified. In both species, GSTO1 mRNA is expressed in liver and heart as well as a range of other tissues. GSTO2 is expressed predominantly in the testis, although moderate levels of expression are seen in other tissues. Extensive immunohistochemistry of rat and human tissue sections has demonstrated cellular and subcellular specificity in the expression of GSTO1-1. The crystal structure of recombinant human GSTO1-1 has been determined, and it adopts the canonical GST fold. A cysteine residue in place of the catalytic tyrosine or serine residues found in other GSTs was shown to form a mixed disulfide with glutathione. Omega class GSTs have dehydroascorbate reductase and thioltransferase activities and also catalyze the reduction of monomethylarsonate, an intermediate in the pathway of arsenic biotransformation. Other diverse actions of human GSTO1-1 include modulation of ryanodine receptors and interaction with cytokine release inhibitory drugs. In addition, GSTO1 has been linked to the age at onset of both Alzheimer's and Parkinson's diseases. Several polymorphisms have been identified in the coding regions of the human GSTO1 and GSTO2 genes. Our laboratory has expressed recombinant human GSTO1-1 and GSTO2-2 proteins, as well as a number of polymorphic variants. The expression and purification of these proteins and determination of their enzymatic activity is described.

Amino Acid Sequence↗

Genome annotation techniques: new approaches and challenges.

As more of the human genome draft sequence is finished, and genomes from other organisms begin to be sequenced, the demand for accurate and reliable genome annotation will increase significantly. To facilitate this industrial-scale genome annotation, automated bioinformatics solutions are increasingly required. As a result, automatic genome annotation systems have become more important in gene discovery within recent years. The design of such large-scale bioinformatics systems is an evolving and dynamic field, based on central cores of bioinformatics software tools and relational databases. Not only must these systems efficiently manage and integrate large volumes of genomic data, but they must also deliver accurate gene predictions and effectively distribute annotation data to the biosciences community.

Computational Biology↗

Molecular and biochemical characterisation of a novel sulphatase gene: Arylsulfatase G (ARSG).

Molecular analysis has provided important insights into the biochemistry and genetics of the sulphatase family of enzymes. Through bioinformatic searches of the EST database, we have identified a novel gene consisting of 11 exons and encoding a 525 aa protein that shares a high degree of sequence similarity with all sulphatases and in particular with arylsulphatases, hence the tentative name Arylsulfatase G (ARSG). The highest homology is shared with Arylsulfatase A, a lysosomal sulphatase which is mutated in metachromatic leukodistrophy, particularly in the amino-terminal region. The 10 amino acids that form the catalytic site are strongly conserved. The murine homologue of Arylsulfatase G gene product shows 87% identity with the human protein. To test the function of this novel gene we transfected the full-length cDNA in Cos7 cells, and detected an Arylsulfatase G precursor protein of 62 kDa. After glycosylation the precursor is maturated in a 70 kDa form, which localises to the endoplasmic reticulum. Northern blot analysis of Arylsulfatase G revealed a ubiquitous expression pattern. We tested the sulphatase activity towards two different artificial substrates 4-methylumbelliferyl (4-MU) sulphate and p-nitrocatechol sulphate, but no arylsulphatase activity was detectable. Further studies are needed to characterise the function of Arylsulfatase G, possibly revealing a novel metabolic pathway.

Amino Acid Sequence↗

A remote and highly conserved enhancer supports amygdala specific expression of the gene encoding the anxiogenic neuropeptide substance-P.

The neuropeptide substance P (SP), encoded by the preprotachykinin-A (PPTA) gene, is expressed in the central and medial amygdaloid nucleus, where it plays a critical role in modulating fear and anxiety related behaviour. Determining the regulatory systems that support PPTA expression in the amygdala may provide important insights into the causes of depression and anxiety related disorders and will provide avenues for the development of novel therapies. In order to identify the tissue specific regulatory element responsible for supporting expression of the PPTA gene in the amygdala, we used long-range comparative genomics in combination with transgenic analysis and immunohistochemistry. By comparing human and chicken genomes, it was possible to detect and characterise a highly conserved long-range enhancer that supported tissue specific expression in SP expressing cells of the medial and central amygdaloid bodies (ECR1; 158.5 kb 5' of human PPTA ORF). Further bioinformatic analysis using the TRANSFAC database indicated that the ECR1 element contained multiple and highly conserved consensus binding sequences of transcription factors (TFs) such as MEIS1. The results of immunohistochemical analysis of transgenic lines were consistent with the hypothesis that the MEIS1 TF interacts with and maintains ECR1 activity in the central amygdala in vivo. The discovery of ECR1 and the in vivo functional relationship with MEIS1 inferred by our studies suggests a mechanism to the regulatory systems that control PPTA expression in the amygdala. Uncovering these mechanisms may play an important role in the future development of tissue specific therapies for the treatment of anxiety and depression.

Amygdala↗

Identification of a novel cell cycle regulated gene, HURP, overexpressed in human hepatocellular carcinoma.

An analytic strategy was followed to identify putative regulatory genes during the development of human hepatocellular carcinoma (HCC). This strategy employed a bioinformatics analysis that used a database search to identify genes, which are differentially expressed in human HCC and are also under cell cycle regulation. A novel cell cycle regulated gene (HURP) that is overexpressed in HCC was identified. Full-length cDNAs encoding the human and mouse HURP genes were isolated. They share 72 and 61% identity at the nucleotide level and amino-acid level, respectively. Endogenous levels of HURP mRNA were found to be tightly regulated during cell cycle progression as illustrated by its elevated expression in the G(2)/M phase of synchronized HeLa cells and in regenerating mouse liver after partial hepatectomy. Immunofluorescence studies revealed that hepatoma up-regulated protein (HURP) localizes to the spindle poles during mitosis. Overexpression of HURP in 293T cells resulted in an enhanced cell growth at low serum levels and at polyhema-based, anchorage-independent growth assay. Taken together, these results strongly suggest that HURP is a potential novel cell cycle regulator that may play a role in the carcinogenesis of human cancer cells.

Amino Acid Sequence↗

Is the Rehydrin TrDr3 from Tortula ruralis associated with tolerance to cold, salinity, and reduced pH? Physiological evaluation of the TrDr3-orthologue, HdeD from Escherichia coli in response to abiotic stress.

We have employed EST analysis in the resurrection moss Tortula ruralis to discover genes that control vegetative desiccation tolerance and describe the characterization of the EST-derived cDNA TrDr3 ( Tortula ruralis desiccation-stress related). The deduced polypeptide TRDR3 has a predicted molecular mass of 25.5 kDa, predicted pI of 6.7, and six transmembrane helical domains. Preliminary expression analyses demonstrate that the TrDr3 transcript ratio increases in response to slow desiccation relative to the hydrated control in both total and polysomal mRNA (mRNP fraction), which classifies TrDr3 as a rehydrin. Bioinformatic searches of the electronic databases reveal that Tortula TRDR3 shares significant similarities to the hdeD gene product ( HNS- dependent expression) from Escherichia coli. The function of the HdeD protein in E. coli is unknown, but it is postulated to be involved in a mechanism of acid stress defence. To establish the role of E. coli HdeD in abiotic stress tolerance, we determined the log survival percentage from shaking cultures of wild-type bacteria and the isogenic hdeD deletion strain (Delta hdeD) in the presence of low temperature (28 degrees C), elevated NaCl (5 % (w/v)), or decreased pH (4.5), or all treatments simultaneously. The Delta hdeD deletion strain was less sensitive, as compared to wild-type E. coli, in response to decreased pH ( p > 0.009), and the combination of all three stresses ( p > 0.0001).

Acclimatization↗

Genetic and molecular markers of urothelial premalignancy and malignancy.

The molecular genetic changes reported in bladder tumors can be classified as primary and secondary aberrations. Primary molecular alterations may be defined as those directly related to the genesis of cancer. These are frequently found as the sole abnormality and are often associated with particular tumors. There are characteristic primary abnormalities involved in th production of low-grade/well-differentiated neoplasms, which destabilize cellular proliferation but have little effect on cellula "social" interactions or differentiation, as well as the rate of cell death or apoptosis. Other molecular events lead to high-grad neoplasms which disrupt growth control, including the cell cycle and apoptosis, and which have a major impact on biological behavior. A primary target leading to low-grade papillary superficial bladder tumors resides on chromosome 9, while p53 gene alterations are commonly seen in flat carcinoma in situ. Other molecular alterations must be elucidated, as many non-invasive neoplasms have neither chromosome 9 nor p53 alterations. Novel approaches utilizing tissue microdissection techniques an molecular genetic assays are needed to shed further light on this subject. Secondary genetic or epigenetic abnormalities may be fortuitous, or may determine the biological behavior of the tumor. Multiple molecular abnormalities are identified in most human cancers studied, including bladder neoplasms. The accumulation, rather than the order, of these genetic alterations may be the critical factor that grants synergistic activity. In this regard, it is noteworthy that many of the genes that are altered act upon the two recognized critical growth and senescenc pathways, TP53 and RB. These particular molecular aberrations may be especially important to evaluate for their use in the management of bladder cancer because of their commonality in progressive forms of the disease. Thus, clinical trials are underway to explore their use in specific situations, particularly in the surgical management of locally advanced disease, and to determine whether adjuvant chemotherapy in such patients may be of benefit. The use of molecular alterations in the management of non-invasive bladder neoplasms remains to be firmly established. Our knowledge of molecular alterations important in bladder cancer progression is far from complete, and further study is necessary to further elucidate cruci pathways involved in progression and therapeutic response. As per preneoplastic conditions, difficulties in identifying and interpreting the significance of phenotypic changes have imposed certain limitations, as has an evolving nomenclature and issues of reproducibility in interpreting morphologica criteria. Nevertheless, molecular alterations involving chromosome 9q and the INK4A locus in papillary superficial tumors vs changes in chromosomes 14q and 8q, p53 and RB in flat carcinoma in situ lesions may indicate a molecular basis for early events that lead to varying pathways in urothelial tumorigenesis. Studies aimed at revealing the clinical relevance of genet instability, as well as molecular or epigenetic alterations, in urothelium and preneoplastic lesions of otherwise morphologicall normal appearance are needed to further advance knowledge in the field. Clinical advances in bladder cancer will be facilitated by novel animal models paralleling the human disease. Molecular diagnostics, particularly specific antigen expression, fluorescence in situ hybridization and microsatellite analyses, have show great promise as screening and follow-up methodologies, and may supplement urine cytology in the diagnosis and characterization of new and recurrent disease. In addition, the use of high-throughput genomic/proteomic assays, linked to comprehensive databases, and coupled with robust bioinformatics will be key elements in elucidating the components of regulatory and signaling pathways involved in bladder tumorigenesis and cancer progression.

Carcinoma in Situ↗

Domain analysis of fatty acid synthase protein (NP_217040) from Mycobacterium tuberculosis H37Rv--a bioinformatics study.

Different domains of fatty acid synthase (FAS) protein of Mycobacterium tuberculosis H37Rv, involved in mycolic acid synthesis were analyzed using various bioinformatics tools. Based on different database searches (CDD and Pfam), FAS protein of Mycobacterium tuberculosis was grouped into eight domains, five of which showed close similarity with pdb templates (1MLA, 1IQ6A, 2BMOA, and 1J3NA). Based on the PSI blast analysis, 3D structures of only five domains were predicted using MODELLER software, and loop modeling was done for only those regions that were predicted as loops by predict protein server. Compared to the original structure, the loop modeled structure showed a lower DOPE score value for FAS protein. The X-ray determined templates that were used for predicting the 3D structure suggest that, FAS protein has "Malonyl-coenzyme A-Hydratase-Nitrobenzene dioxygenase-3-oxoacyl-(acp) synthase" activity. Accuracy of the prediction of 3D structure of different domains of FAS protein was further validated by Ramachandran plot and PROCHECK (G-value).

Amino Acid Sequence↗