Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

LEGER: knowledge database and visualization tool for comparative genomics of pathogenic and non-pathogenic Listeria species.

Listeria species are ubiquitous in the environment and often contaminate foods because they grow under conditions used for food preservation. Listeria monocytogenes, the human and animal pathogen, causes Listeriosis, an infection with a high mortality rate in risk groups such as immune-compromised individuals. Furthermore, L.monocytogenes is a model organism for the study of intracellular bacterial pathogens. The publication of its genome sequence and that of the non-pathogenic species Listeria innocua initiated numerous comparative studies and efforts to sequence all species comprising the genus. The Proteome database LEGER (http://leger2.gbf.de/cgi-bin/expLeger.pl) was developed to support functional genome analyses by combining information obtained by applying bioinformatics methods and from public databases to improve the original annotations. LEGER offers three unique key features: (i) it is the first comprehensive information system focusing on the functional assignment of genes and proteins; (ii) integrated visualization tools, KEGG pathway and Genome Viewer, alleviate the functional exploration of complex data; and (iii) LEGER presents results of systematic post-genome studies, thus facilitating analyses combining computational and experimental results. Moreover, LEGER provides an unpublished membrane proteome analysis of L.innocua and in total visualizes experimentally validated information about the subcellular localizations of 789 different listerial proteins.

Bacterial Proteins↗

Implications of structural genomics target selection strategies: Pfam5000, whole genome, and random approaches.

Structural genomics is an international effort to determine the three-dimensional shapes of all important biological macromolecules, with a primary focus on proteins. Target proteins should be selected according to a strategy that is medically and biologically relevant, of good value, and tractable. As an option to consider, we present the "Pfam5000" strategy, which involves selecting the 5000 most important families from the Pfam database as sources for targets. We compare the Pfam5000 strategy to several other proposed strategies that would require similar numbers of targets. These strategies include complete solution of several small to moderately sized bacterial proteomes, partial coverage of the human proteome, and random selection of approximately 5000 targets from sequenced genomes. We measure the impact that successful implementation of these strategies would have upon structural interpretation of the proteins in Swiss-Prot, TrEMBL, and 131 complete proteomes (including 10 of eukaryotes) from the Proteome Analysis database at the European Bioinformatics Institute (EBI). Solving the structures of proteins from the 5000 largest Pfam families would allow accurate fold assignment for approximately 68% of all prokaryotic proteins (covering 59% of residues) and 61% of eukaryotic proteins (40% of residues). More fine-grained coverage that would allow accurate modeling of these proteins would require an order of magnitude more targets. The Pfam5000 strategy may be modified in several ways, for example, to focus on larger families, bacterial sequences, or eukaryotic sequences; as long as secondary consideration is given to large families within Pfam, coverage results vary only slightly. In contrast, focusing structural genomics on a single tractable genome would have only a limited impact in structural knowledge of other proteomes: A significant fraction (about 30-40% of the proteins and 40-60% of the residues) of each proteome is classified in small families, which may have little overlap with other species of interest. Random selection of targets from one or more genomes is similar to the Pfam5000 strategy in that proteins from larger families are more likely to be chosen, but substantial effort would be spent on small families.

Animals↗

The Ensembl genome database project.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of the human genome sequence, with confirmed gene predictions that have been integrated with external data sources, and is available as either an interactive web site or as flat files. It is also an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements from sequence analysis to data storage and visualisation. The Ensembl site is one of the leading sources of human genome sequence annotation and provided much of the analysis for publication by the international human genome project of the draft genome. The Ensembl system is being installed around the world in both companies and academic sites on machines ranging from supercomputers to laptops.

Computational Biology↗

Ensembl 2002: accommodating comparative genomics.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of human, mouse and other genome sequences, available as either an interactive web site or as flat files. Ensembl also integrates manually annotated gene structures from external sources where available. As well as being one of the leading sources of genome annotation, Ensembl is an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements. These range from sequence analysis to data storage and visualisation and installations exist around the world in both companies and at academic sites. With both human and mouse genome sequences available and more vertebrate sequences to follow, many of the recent developments in Ensembl have focusing on developing automatic comparative genome analysis and visualisation.

Animals↗

Ensembl 2004.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organize biology around the sequences of large genomes. It is a comprehensive and integrated source of annotation of large genome sequences, available via interactive website, web services or flat files. As well as being one of the leading sources of genome annotation, Ensembl is an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements. The facilities of the system range from sequence analysis to data storage and visualization and installations exist around the world both in companies and at academic sites. With a total of nine genome sequences available from Ensembl and more genomes to follow, recent developments have focused mainly on closer integration between genomes and external data.

Animals↗

Developmental mouse brain gene expression maps.

Brain gene expression databases are providing an increasing amount of information to the neuroscience community. Most databases are focused on the adult mouse rather than embryonic development. Here we survey the major mouse gene expression databases for the developing brain. The high throughput in situ hybridization approach generates large volumes of gene expression data that can be compiled and examined in a relatively short period of time. It is of increasing importance to compare gene expression patterns of neurodevelopment in the brain in relation to the adult. Often clues to adult gene expression and gene function can be determined by examining embryonic development. It is our hope that once all genes are mapped in the brain from the embryo to the adult, studies can be conducting based on information derived from such databases in conjunction with other bioinformatics sources.

Animals↗

PROFILER: a tool for automatic searching of internally maintained databases.

A new application program, the BIOINFORMATICS PROFILER, is described which simplifies the analysis of new genome sequence information appearing on a daily basis. Control tasks are defined through an intuitive graphical user interface and are executed at user defined nightly intervals. Electronic mail is sent to indicate that search results have been found. All task output is presented to users in the form of hypertext (HTML), allowing easy browsing. Currently supported tasks include BLAST and FastA sequence searching, keyword based searching of network news articles and WAIS databases, examination of GenBank sequence entries using regular expressions and boolean operations and protein sequence motif searching.

Amino Acid Sequence↗

Bioinformatic assessment of mass spectrometric chemical derivatisation techniques for proteome database searching.

Identification of proteins from the mass spectra of peptide fragments generated by proteolytic cleavage using database searching has become one of the most powerful techniques in proteome science, capable of rapid and efficient protein identification. Using computer simulation, we have studied how the application of chemical derivatisation techniques may improve the efficiency of protein identification from mass spectrometric data. These approaches enhance ion yield and lead to the promotion of specific ions and fragments, yielding additional database search information. The impact of three alternative techniques has been assessed by searching representative proteome databases for both single proteins and simple protein mixtures. For example, by reliably promoting fragmentation of singly-charged peptide ions at aspartic acid residues after homoarginine derivatisation, 82% of yeast proteins can be unambiguously identified from a single typical peptide-mass datum, with a measured mass accuracy of 50 ppm, by using the associated secondary ion data. The extra search information also provides a means to confidently identify proteins in protein mixtures where only limited data are available. Furthermore, the inclusion of limited sequence information for the peptides can compensate and exceed the search efficiency available via high accuracy searches of around 5 ppm, suggesting that this is a potentially useful approach for simple protein mixtures routinely obtained from two-dimensional gels.

Animals↗

Intregrated analysis of the human cardiac transcriptome, proteome and phosphoproteome.

Altered expression of different classes of genes has been shown to differentiate between failing and nonfailing human hearts. However, characterization of proteins and the post-translational modifications that regulate their functions is required for understanding both the physiology of cardiac muscle and the mechanisms leading to pathological states associated with cardiac diseases. We present in this paper, an analysis of the human cardiac transcriptome, proteome and phosphoproteome. Data from two sources (i) experiments performed in our laboratory and (ii) bioinformatics searches of public databases (SWISS-PROT, NCBI, Cardiac Gene Expression Knowledge Base, Gene Ontology Consortium and Affymetrix) are reported in a relational database that allows user-designed specific queries. Microarray experiments were performed with Affymetrix Hu95Av2. Cardiac proteins were digested with trypsin. An 11 step cation exchange procedure produced fractions for analysis in separate reversed phase high-performance liquid chromatography-tandem mass spectrometry (MS/MS) experiments. Immobilized metal affinity chromatography was used to select the phosphopeptides from the same tryptic peptide mixture. They were then further investigated by MS/MS. Gel-free approaches were used to detect 267 proteins and 47 phosphopeptides. Our human cardiac database contains 447 entries. We propose the use of this platform, built with data derived from nonfailing hearts, as a template for initiating the effort to characterize the human cardiac proteome and its associated post-translational modifications.

Chromatography, Affinity↗

Ethnopharmacology and bioinformatic combination for leads discovery: application to phospholipase A(2) inhibitors.

A program combining ethnopharmacology and bioinformatic approaches has successfully been applied on anti-inflammatory activity. (i) An ethnobotanical study allowed the identification of several plants associated with putative anti-inflammatory properties as potential leads. (ii) On the other hand, it is well known that phospholipase A(2) is a target implicated in the pro-inflammatory process. Thus, (iii) some selected plant extracts were experimentally tested on phospholipase A(2). Finally, (iv) these experimental results combined with bioinformatic tools, such as database exploitation and molecular modeling, allowed to suggest that one compound, betulin and its oxidative form betulinic acid, might be responsible of the anti-PLA(2) activity. This suggestion was confirmed experimentally.

Chromatography, High Pressure Liquid↗

The Southeast Collaboratory for Structural Genomics: a high-throughput gene to structure factory.

The Southeast Collaboratory for Structural Genomics consists of four working groups. The protein production group supplies/develops high-output production of Pyrococcus furiosus, Caenorhabditis elegans, and selected human proteins. The X-ray crystallography group conducts high-throughput structure production in parallel with production-related research/development in nanocrystallization robotics, capillary crystallization cassette, synchrotron/home X-ray instrumentation, sample mounting robotics, data processing and pipelined structure analysis, combined refinement/validation protocols, and direct use of unlabeled native crystals (Direct Crystallography). The NMR group emphasizes/develops sample screening and backbone structure determination from residual dipolar coupling data. The bioinformatics group implements/develops local database interfaces, pipelined sequence/structure information search/updates, and database/bioinformatics toolkits.

Animals↗

ANXA3 hypomethylation as a prognostic biomarker in hepatitis B virus-related acute-on-chronic liver failure.

BACKGROUND: Hepatitis B virus-related acute-on-chronic liver failure (HBV-ACLF) is associated with a poor prognosis. This research aimed to characterize the expression pattern and clinical value of Annexin A3 (ANXA3) in HBV-ACLF patients. METHODS: First of all, ACLF-related datasets were downloaded from the Gene Expression Omnibus (GEO) database to carry out bioinformatics analyses. RT-qPCR, ELISA, and Methylight were used to measure ANXA3 gene expression and promoter methylation levels. A validation cohort was leveraged to further validate the results. RESULTS: Transcriptome analysis showed that ANXA3 was among the most differentially expressed genes when comparing dead patients with HBV-ACLF to those with survivors. The mRNA and serum levels of ANXA3 were elevated, and methylation levels were decreased in HBV-ACLF patients. The PMR value of ANXA3 in patients with HBV-ACLF was negatively correlated with inflammation-related cytokines IL-6, TNF-&#x3b1;, and IL-1&#x3b2;, as well as quantitative clinical parameters AST, TBIL, PT, INR, NEUT%, and MELD score, and positively correlated with PTA (all p&#x2009;<&#x2009;0.05). In HBV-ACLF patients, ANXA3 was considered to be an independent influence factor for the 90-day mortality. It was also found that ANXA3, especially hypomethylation, was associated with 28- and 90-day overall survival in patients with HBV-ACLF based on receiver operating characteristic (ROC) analysis, decision curve analysis (DCA), and Kaplan-Meier curves. CONCLUSIONS: ANXA3 hypomethylation has a prominent predictive value for short-term mortality in patients with HBV-ACLF and may serve as a promising biomarker of HBV-ACLF prognosis.

Humans↗

Evolutionary analysis of fructose 2,6-bisphosphate metabolism.

Fructose 2,6-bisphosphate is a potent metabolic regulator in eukaryotic organisms; it affects the activity of key enzymes of the glycolytic and gluconeogenic pathways. The enzymes responsible for its synthesis and hydrolysis, 6-phosphofructo-2-kinase (PFK-2) and fructose-2,6-bisphosphatase (FBPase-2) are present in representatives of all major eukaryotic taxa. Results from a bioinformatics analysis of genome databases suggest that very early in evolution, in a common ancestor of all extant eukaryotes, distinct genes encoding PFK-2 and FBPase-2, or related enzymes with broader substrate specificity, fused resulting in a bifunctional enzyme both domains of which had, or later acquired, specificity for fructose 2,6-bisphosphate. Subsequently, in different phylogenetic lineages duplications of the gene of the bifunctional enzyme occurred, allowing the development of distinct isoenzymes for expression in different tissues, at specific developmental stages or under different nutritional conditions. Independently in different lineages of many unicellular eukaryotes one of the domains of the different PFK-2/FBPase-2 isoforms has undergone substitutions of critical catalytic residues, or deletions rendering some enzymes monofunctional. In a considerable number of other unicellular eukaryotes, mainly parasitic organisms, the enzyme seems to have been lost altogether. Besides the catalytic core, the PFK-2/FBPase-2 has often N- and C-terminal extensions which show little sequence conservation. The N-terminal extension in particular can vary considerably in length, and seems to have acquired motifs which, in a lineage-specific manner, may be responsible for regulation of catalytic activities, by phosphorylation or ligand binding, or for mediating protein-protein interactions.

Animals↗

GARSA: genomic analysis resources for sequence annotation.

SUMMARY: Growth of genome data and analysis possibilities have brought new levels of difficulty for scientists to understand, integrate and deal with all this ever-increasing information. In this scenario, GARSA has been conceived aiming to facilitate the tasks of integrating, analyzing and presenting genomic information from several bioinformatics tools and genomic databases, in a flexible way. GARSA is a user-friendly web-based system designed to analyze genomic data in the context of a pipeline. EST and GGS data can be analyzed using the system since it accepts (1) chromatograms, (2) download of sequences from GenBank, (3) Fasta files stored locally or (4) a combination of all three. Quality evaluation of chromatograms, vector removing and clusterization are easily performed as part of the pipeline. A number of local and customizable Blast and CDD analyses can be performed as well as Interpro, complemented with phylogeny analyses. GARSA is being used for the analyses of Trypanosoma vivax (GSS and EST), Trypanosoma rangeli (GSS, EST and ORESTES), Bothrops jararaca (EST), Piaractus mesopotamicus (EST) and Lutzomyia longipalpis (EST). AVAILABILITY: The GARSA system is freely available under GPL license (http://www.biowebdb.org/garsa/). For download requests visit http://www.biowebdb.org/garsa/ or contact Dr Alberto Dávila.

Animals↗

A reference database for circular dichroism spectroscopy covering fold and secondary structure space.

MOTIVATION: Circular Dichroism (CD) spectroscopy is a long-established technique for studying protein secondary structures in solution. Empirical analyses of CD data rely on the availability of reference datasets comprised of far-UV CD spectra of proteins whose crystal structures have been determined. This article reports on the creation of a new reference dataset which effectively covers both secondary structure and fold space, and uses the higher information content available in synchrotron radiation circular dichroism (SRCD) spectra to more accurately predict secondary structure than has been possible with existing reference datasets. It also examines the effects of wavelength range, structural redundancy and different means of categorizing secondary structures on the accuracy of the analyses. In addition, it describes a novel use of hierarchical cluster analyses to identify protein relatedness based on spectral properties alone. The databases are shown to be applicable in both conventional CD and SRCD spectroscopic analyses of proteins. Hence, by combining new bioinformatics and biophysical methods, a database has been produced that should have wide applicability as a tool for structural molecular biology.

Algorithms↗

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software↗

The MIGenAS integrated bioinformatics toolkit for web-based sequence analysis.

We describe a versatile and extensible integrated bioinformatics toolkit for the analysis of biological sequences over the Internet. The web portal offers convenient interactive access to a growing pool of chainable bioinformatics software tools and databases that are centrally installed and maintained by the RZG. Currently, supported tasks comprise sequence similarity searches in public or user-supplied databases, computation and validation of multiple sequence alignments, phylogenetic analysis and protein-structure prediction. Individual tools can be seamlessly chained into pipelines allowing the user to conveniently process complex workflows without the necessity to take care of any format conversions or tedious parsing of intermediate results. The toolkit is part of the Max-Planck Integrated Gene Analysis System (MIGenAS) of the Max Planck Society available at www.migenas.org (click 'Start Toolkit').

Animals↗

A novel start-loss mutation of the SLC29A3 gene in a consanguineous family with H syndrome: clinical characteristics, in silico analysis and literature review.

BACKGROUND: The SLC29A3 gene, which encodes a nucleoside transporter protein, is primarily located in intracellular membranes. The mutations in this gene can give rise to various clinical manifestations, including H syndrome, dysosteosclerosis, Faisalabad histiocytosis, and pigmented hypertrichosis with insulin-dependent diabetes. The aim of this study is to present two Iranian patients with H syndrome and to describe a novel start-loss mutation in SLC29A3 gene. METHODS: In this study, we employed whole-exome sequencing (WES) as a method to identify genetic variations that contribute to the development of H syndrome in a 16-year-old girl and her 8-year-old brother. These siblings were part of an Iranian family with consanguineous parents. To confirmed the pathogenicity of the identified variant, we utilized in-silico tools and cross-referenced various databases to confirm its novelty. Additionally, we conducted a co-segregation study and verified the presence of the variant in the parents of the affected patients through Sanger sequencing. RESULTS: In our study, we identified a novel start-loss mutation (c.2T&#x2009;>&#x2009;A, p.Met1Lys) in the SLC29A3 gene, which was found in both of two patients. Co-segregation analysis using Sanger sequencing confirmed that this variant was inherited from the parents. To evaluate the potential pathogenicity and novelty of this mutation, we consulted various databases. Additionally, we employed bioinformatics tools to predict the three-dimensional structure of the mutant SLC29A3 protein. These analyses were conducted with the aim of providing valuable insights into the functional implications of the identified mutation on the structure and function of the SLC29A3 protein. CONCLUSION: Our study contributes to the expanding body of evidence supporting the association between mutations in the SLC29A3 gene and H syndrome. The molecular analysis of diseases related to SLC29A3 is crucial in understanding the range of variability and raising awareness of H syndrome, with the ultimate goal of facilitating early diagnosis and appropriate treatment. The discovery of this novel biallelic variant in the probands further underscores the significance of utilizing genetic testing approaches, such as WES, as dependable diagnostic tools for individuals with this particular condition.

Humans↗