Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Combining bioinformatics resources for the structural modelling of eukaryotic metabolic networks.

The architecture of the cellular metabolic network is almost completely available from several databases. This has paved the way for computational analyses. Whereas kinetic modelling is still restrained to small metabolic sub-systems for which enzyme-kinetic details are known, so-called structural modelling techniques can be applied to complete metabolic networks even if the kinetics and regulation of the underlying enzymes is still unknown. Structural modelling requires detailed information on the presence of metabolic enzymes in a specific cell type of interest and the thermodynamics of the reactions, determining their direction under cellular conditions. If compartments are distinguished the sub-cellular compartmentation of reactions and enzymes and the membrane transporters exchanging metabolites between cellular compartments must be included. All this information cannot be taken from a single data base but has to be compiled from various Bioinformatics resources. Here we present an approach towards the organization of Bioinformatics data that enables the flux-balance analysis of comprehensive compartmentalized metabolic networks of eukaryotic cells with special focus on human hepatocytes.

Aspartate Aminotransferases↗

Testing chromosomal phylogenies and inversion breakpoint reuse in Drosophila.

A combination of cytogenetic and bioinformatic procedures was used to test the chromosomal phylogeny relating Drosophila buzzatii with D. repleta. Chromosomes X and 2, harboring most of the inversions fixed between these two species, were analyzed. First, chromosomal segments conserved during the divergence of the two species were identified by comparative in situ hybridization to the D. repleta chromosomes of 180 BAC clones from a BAC-based physical map of the D. buzzatii genome. These conserved segments were precisely delimited with the aid of clones containing inversion breakpoints. Then GRIMM software was used to estimate the minimum number of rearrangements necessary to transform one genome into the other and identify all possible rearrangement scenarios. Finally, the most plausible inversion trajectory was tested by hybridizing 12 breakpoint-bearing BAC clones to the chromosomes of seven other species in the repleta group. The results show that chromosomes X and 2 of D. buzzatii and D. repleta differ by 12 paracentric inversions. Nine of them are fixed in chromosome 2 and entail two breakpoint reuses. Our results also show that the cytological relationship between D. repleta and D. mercatorum is closer than that between D. repleta and D. peninsularis, and we propose that the phylogenetic relationships in this lineage of the repleta group be reconsidered. We also estimated the rate of rearrangement between D. repleta and D. buzzatii and conclude that rates within the genus Drosophila vary substantially between lineages, even within a single species group.

Animals↗

Formalization of mouse embryo anatomy.

MOTIVATION: The Edinburgh Mouse Atlas and Gene Expression Database project has developed a digital atlas of mouse development to provide a spatio-temporal framework for spatially mapped data such as in situ gene expression and cell lineage. As part of this database, a mouse embryo anatomy ontology has been created. A formalization of this anatomy is required to document its precise semantics and how it is used in the context of the Mouse Atlas. RESULTS: The paper describes the existing anatomy ontology and formalizes aspects of it using a predicate logic based approach. It therefore provides a guide for users of the current version of the ontology, as well as the basis for a description of the anatomy using an ontology language, such as OWL, thus enabling future work on reasoning about the Mouse Atlas in the context of an intelligent gene expression bioinformatics workflow system. The logic has been implemented in a Prolog prototype. AVAILABILITY: The Mouse Atlas is available on-line at http://genex.hgu.mrc.ac.uk

Algorithms↗

A novel splice-altering TNC variant (c.5247A > T, p.Gly1749Gly) in an Chinese family with autosomal dominant non-syndromic hearing loss.

BACKGROUND: This study aims to analyze the pathogenic gene in a Chinese family with non-syndromic hearing loss and identify a novel mutation site in the TNC gene. METHODS: A five-generation Chinese family from Anhui Province, presenting with autosomal dominant non-syndromic hearing loss, was recruited for this study. By analyzing the family history, conducting clinical examinations, and performing genetic analysis, we have thoroughly investigated potential pathogenic factors in this family. The peripheral blood samples were obtained from 20 family members, and the pathogenic genes were identified through whole exome sequencing. Subsequently, the mutation of gene locus was confirmed using Sanger sequencing. The conservation of TNC mutation sites was assessed using Clustal Omega software. We utilized functional prediction software including dbscSNV_AdaBoost, dbscSNV_RandomForest, NNSplice, NetGene2, and Mutation Taster to accurately predict the pathogenicity of these mutations. Furthermore, exon deletions were validated through RT-PCR analysis. RESULTS: The family exhibited autosomal dominant, progressive, post-lingual, non-syndromic hearing loss. A novel synonymous variant (c.5247A > T, p.Gly1749Gly) in TNC was identified in affected members. This variant is situated at the exon-intron junction boundary towards the end of exon 18. Notably, glycine residue at position 1749 is highly conserved across various species. Bioinformatics analysis indicates that this synonymous mutation leads to the disruption of the 5' end donor splicing site in the 18th intron of the TNC gene. Meanwhile, verification experiments have demonstrated that this synonymous mutation disrupts the splicing process of exon 18, leading to complete exon 18 skipping and direct splicing between exons 17 and 19. CONCLUSION: This novel splice-altering variant (c.5247A > T, p.Gly1749Gly) in exon 18 of the TNC gene disrupts normal gene splicing and causes hearing loss among HBD families.

Adult↗

bioTk:componentry for genome informatics graphical user interfaces.

bioTk is a collection of graphical "widgets" and utilities that support application programming in the domain of bioinformatics. It is intended to establish a framework that encourages the development of communicating window-based applications and flexible, non-modal user interaction. The current release of bioTk has domain-specific widgets for chromosome ideogram displays, genome maps, and scrolling sequence windows.

Base Sequence↗

fjoin: simple and efficient computation of feature overlaps.

Sets of biological features with genome coordinates (e.g., genes and promoters) are a particularly common form of data in bioinformatics today. Accordingly, an increasingly important processing step involves comparing coordinates from large sets of features to find overlapping feature pairs. This paper presents fjoin, an efficient, robust, and simple algorithm for finding these pairs, and a downloadable implementation. For typical bioinformatics feature sets, fjoin requires O(n log(n)) time (O(n) if the inputs are sorted) and uses O(1) space. The reference implementation is a stand-alone Python program; it implements the basic algorithm and a number of useful extensions, which are also discussed in this paper.

Algorithms↗

Novel technologies and recent advances in metastasis research.

In this review we have attempted to summarize some of the recent developments in using novel technologies to unravel the molecular mechanisms of tumor progression, in particular the formation of tumor metastasis. In order to push forward the frontiers in cancer research, it is obvious that several fields have to be further developed and interconnected: (1) clinical, epidemiological and pathological studies which mainly use innovative technologies, including microarray technology and nanotechnology to determine as many parameters as possible, (2) the development of improved and suitable bioassays and better animal models and (3) the use of novel computation and bioinformatics methods to sample and integrate the exponentially growing sets of data coming from such investigations. Fashionable as scientists are, this new endeavor may be called systems biology.

Animals↗

Exploring shared biomarkers and their mechanisms in thyroid cancer and systemic lupus erythematosus via bioinformatics analysis.

BACKGROUND: Systemic lupus erythematosus (SLE), an autoimmune disorder, is linked to a heightened risk of multiple malignancies, including thyroid cancer. Thyroid cancer is the most prevalent malignancy of the endocrine system, and its autoimmune-related pathological features render it an optimal subject for investigating the mechanisms of their comorbidity. The molecular mechanisms underlying this comorbidity are still ambiguous. The accurate diagnosis and treatment of thyroid cancer urgently necessitate innovative molecular targets that extend beyond conventional pathological characteristics. This study seeks to employ integrated bioinformatics approaches to elucidate potential shared molecular mechanisms and immunological features between thyroid cancer and systemic lupus erythematosus (SLE), aiming to enhance understanding of their comorbidity and identify novel intervention targets. METHODS: This study initially acquired gene expression data for TC and SLE from the GEO database and subsequently screened and identified differentially expressed genes (DEGs) shared by both diseases. Subsequently, we conducted Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Reactome functional enrichment analyses on these 46 shared differentially expressed genes (DEGs) and further assessed the activation status of pertinent pathways using Gene Set Enrichment Analysis (GSEA). Subsequently, we employed CIBERSORTx to examine immune infiltration patterns and developed protein-protein interaction networks utilising the STRING database. We identified hub genes utilising the MCODE and cytoHubba plugins and visualised the findings with Cytoscape software. We additionally assessed the diagnostic efficacy of these core hub genes in an independent dataset utilising ROC curves and investigated their prognostic relevance in thyroid cancer through Kaplan-Meier survival analysis and multivariate Cox proportional hazards regression. Ultimately, we employed the Network Analyst platform to forecast transcription factor-gene and miRNA-gene regulatory networks and identified potential targeted therapeutic compounds utilising the DSigDB database. RESULTS: This study identified 46 differentially expressed genes (DEGs) commonly linked to thyroid cancer and systemic lupus erythematosus (SLE), which were significantly enriched in signalling pathways associated with immune-inflammatory activation, type I interferon responses, and complement pathway activation. Moreover, GSEA findings validated that immune-inflammatory and autoimmune-related pathways are markedly activated in both conditions. Twelve hub genes were discerned through protein-protein interaction networks. Analysis of immune infiltration indicated that thyroid cancer and systemic lupus erythematosus exhibit a shared characteristic of innate immune dysregulation, marked by the infiltration of myeloid cells (neutrophils, M0/M2 macrophages). Receiver operating characteristic (ROC) curve analysis identified six significant core hub genes with substantial diagnostic value: C1QB, LCN2, C1QC, LTF, VSIG4, and C3AR1. Univariate survival analysis indicated that elevated expression of C1QC and C3AR1 significantly enhances overall survival in thyroid cancer patients; however, multivariate COX regression analysis revealed that their independent prognostic significance necessitates further validation. This study predicted the interaction networks of transcription factors and miRNAs regulating key genes, with LCN2 demonstrating the highest connectivity to miRNAs, and identified candidate therapeutic compounds linked to it. CONCLUSION: This study employed bioinformatics analysis to identify critical shared hub genes and molecular pathways connecting thyroid cancer and systemic lupus erythematosus, offering novel insights into their shared pathogenesis and the advancement of targeted biomarkers and therapeutic strategies.

Bioinformatics analysis↗

The untranslated regions of eukaryotic mRNAs: structure, function, evolution and bioinformatic tools for their analysis.

The crucial role of the non-coding portion of genomes is now widely acknowledged. In particular, mRNA untranslated regions are involved in many post-transcriptional regulatory pathways that control mRNA localisation, stability and translation efficiency. A review is given of the most recent research works on the functional characterisation of eukaryotic mRNA untranslated regions. In order to make possible a systematic and detailed sequence analysis of mRNA untranslated regions (UTRs), a non-redundant database of metazoan mRNA untranslated sequences annotated for the occurrence of specific functional elements, UTRdb, was devised. These elements, whose consensus structure has been devised on the basis of experimental assays and of comparative analyses, have been collected in the UTRsite database. A suitable pattern-matching software has been devised to search UTRsite patterns in user-submitted sequences, also assessing their statistical significance. Structural, compositional and evolutionary features of untranslated sequences of metazoan mRNAs have been investigated showing peculiar intra- and interspecific patterns.

Animals↗

AutoFACT: an automatic functional annotation and classification tool.

BACKGROUND: Assignment of function to new molecular sequence data is an essential step in genomics projects. The usual process involves similarity searches of a given sequence against one or more databases, an arduous process for large datasets. RESULTS: We present AutoFACT, a fully automated and customizable annotation tool that assigns biologically informative functions to a sequence. Key features of this tool are that it (1) analyzes nucleotide and protein sequence data; (2) determines the most informative functional description by combining multiple BLAST reports from several user-selected databases; (3) assigns putative metabolic pathways, functional classes, enzyme classes, GeneOntology terms and locus names; and (4) generates output in HTML, text and GFF formats for the user's convenience. We have compared AutoFACT to four well-established annotation pipelines. The error rate of functional annotation is estimated to be only between 1-2%. Comparison of AutoFACT to the traditional top-BLAST-hit annotation method shows that our procedure increases the number of functionally informative annotations by approximately 50%. CONCLUSION: AutoFACT will serve as a useful annotation tool for smaller sequencing groups lacking dedicated bioinformatics staff. It is implemented in PERL and runs on LINUX/UNIX platforms. AutoFACT is available at http://megasun.bch.umontreal.ca/Software/AutoFACT.htm.

Acanthamoeba castellanii↗

Normal/disease-paired cDNA library subtraction for molecular marker development.

This study consists of a novel strategy for the identification of potential cancer exclusively expressed genes, which might lead to the development of valuable diagnostic molecular markers. Normal- and cancer-paired tissues from the same patient were collected and subjected to the construction of primary complementary deoxyribonucleic acid (cDNA) libraries. The cancer cDNA library was generated by subtracting normal cDNAs from the primary cancer library. The remaining clones in the subtracted cancer library were sequenced and Basic Local Alignment Search Tool (BLAST)-analyzed against current nonredundant and est_human databases in the GenBank. The clones that had no matches with any known gene sequences except the human genomic sequences were identified as ideal candidates for potential molecular marker development. The candidate marker sequence and associated primer sequences were identified on the target clone sequence of the region with the least, or no, matched homologs. The specificity of the markers was measured by a polymerase chain reaction test experiment with the DNA templates from different human normal tissues, genomic DNA, and additional patients' tissues. Results from bioinformatical and experimental approaches used suggest that the methodology has a high potential to identify cancer exclusive transcripts. Thus, it might result in the development of diagnostic markers.

Biotechnology↗

PSST-2.0: Protein Data Bank Sequence Search Tool.

UNLABELLED: PSST-2.0 (Protein Data Bank [PDB] Sequence Search Tool) is an updated version of the earlier PSST (Protein Sequence Search Tool), and the philosophy behind the search engine has remained unchanged. PSST-2.0 is a Web-based, interactive search engine developed to retrieve required protein or nucleic acid sequence information and some of its related details, primarily from sequences derived from the structures deposited in the PDB (the database of 3-dimensional [3-D] protein and nucleic acid structures). Additionally, the search engine works for a selected subset of 25% or 90% non-homologous protein chains. For some of the selected options, the search engine produces a detailed output for the user-uploaded, 3-D atomic coordinates of the protein structure (PDB file format) from the client machine through the Web browser. The search engine works on a locally maintained PDB, which is updated every week from the parent server at the Research Collaboratory for Structural Bioinformatics, and hence the search results are up to date at any given time. AVAILABILITY: PSST-2.0 is freely accessible via http://pranag.physics.iisc.ernet.in/psst/ or http://144.16.71.10/psst/.

Algorithms↗

From XML to RDF: how semantic web technologies will change the design of 'omic' standards.

With the ongoing rapid increase in both volume and diversity of 'omic' data (genomics, transcriptomics, proteomics, and others), the development and adoption of data standards is of paramount importance to realize the promise of systems biology. A recent trend in data standard development has been to use extensible markup language (XML) as the preferred mechanism to define data representations. But as illustrated here with a few examples from proteomics data, the syntactic and document-centric XML cannot achieve the level of interoperability required by the highly dynamic and integrated bioinformatics applications. In the present article, we discuss why semantic web technologies, as recommended by the World Wide Web consortium (W3C), expand current data standard technology for biological data representation and management.

Algorithms↗

The Merck Gene Index browser: an extensible data integration system for gene finding, gene characterization and EST data mining.

MOTIVATION: To make effective use of the vast amounts of expressed sequence tag (EST) sequence data generated by the Merck-sponsored EST project and other similar efforts, sequences must be organized into gene classes, and scientists must be able to 'mine' the gene class data in the context of related genomic data. RESULTS: This paper presents the Merck Gene Index browser, an easily extensible, World Wide Web-based system for mining the Merck Gene Index (MGI) and related genomic data. The MGI is a non-redundant set of clones and sequences, each representing a distinct gene, constructed from all high-quality 3' EST sequences generated by the Merck-sponsored EST project. The MGI browser integrates data from a variety of sources and storage formats, both local and remote, using an eclectic integration strategy, including a federation of relational databases, a local data warehouse and simple hypertext links. Data currently integrated include: LENS cDNA clone and EST data, dbEST protein and non-EST nucleic acid similarity data, WashU sequence chromatograms. Entrez sequence and Medline entries, and UniGene gene clusters. Flatfile sequence data are accessed using the Bioapps server, an internally developed client-server system that supports generic sequence analysis applications. Browser data are retrieved and formatted by means of the Bioinformatics Data Integration Toolkit (B-DIT), a new suite of Perl routines.

Abstracting and Indexing↗

Annotated proteome of a human T-cell lymphoma.

As the reliable identification of proteins by tandem mass spectrometry becomes increasingly common, the full characterization of large data sets of proteins remains a difficult challenge. Our goal was to survey the proteome of a human T-cell lymphoma-derived cell line in a single set of experiments and present an automated method for the annotation of lists of proteins. A downstream application of these data includes the identification of novel pathogenetic and candidate diagnostic markers of T-cell lymphoma. Total protein isolated from cytoplasmic, membrane, and nuclear fractions of the SUDHL-1 T-cell lymphoma cell line was resolved by SDS-PAGE, and the entire gel lanes digested and analyzed by tandem mass spectrometry. Acquired data files were searched against the UniProt protein database using the SEQUEST algorithm. Search results for each subcellular fraction were analyzed using INTERACT and ProteinProphet. All protein identifications with an error rate of less than 10% were directly exported into excel and analyzed using GOMiner (NIH/NCI). The Gene ontology molecular function and cell location data were summarized for the identified proteins and results exported as user-interactive directed acyclic graphs. A total of 1105 unique proteins were identified and fully annotated, including numerous proteins that had not been previously characterized in lymphoma, in functional categories such as cell adhesion, migration, signaling, and stress response. This study demonstrates the utility of currently available bioinformatics tools for the robust identification and annotation of large numbers of proteins in a batchwise fashion.

Algorithms↗

In search of candidate genes critically expressed in the human endometrium during the window of implantation.

BACKGROUND: In this prospective randomized blinded clinical trial, we examined gene expression profiles of the human endometrium during the early and mid-luteal phases of the natural cycle. METHODS: An endometrial biopsy was performed on day 16 (LH +3) or on day 21 (LH +8), followed by RNA extraction and microarray analysis using an Affymetrix HG-U95A microchip. Data analysis was carried out using pairwise multiple group comparison with the significance analysis of microarrays (SAM) software. RESULTS: With a false discovery rate of 0, the analysis revealed that 107 genes were significantly and differently expressed (> or =2-fold) during the early versus the mid-luteal phase of the cycle. Forty-five of these genes have not been previously linked to endometrial receptivity. Validation of the microarray data was accomplished using semiquantitative RT-PCR. We demonstrated the presence of estrogen and progesterone response elements (ERE and PRE) by analysis of the 5'-flanking regions of a subset of differentially regulated genes. CONCLUSIONS: Using a strict bioinformatics approach of microarray data, we demonstrated significant changes in candidate genes during the transition of the early to the mid-luteal phase of the human endometrium that may have functional significance for the opening and maintenance of the window of implantation.

Adult↗

Representing bioinformatics causality.

This paper reviews a variety of different graphical notations currently in active use for modelling dynamic processes in bioinformatics and biotechnology, and crystallises from these notations a set of properties essential to any proposal for a modelling language seeking to provide an adequate systemic description of biological processes.

Algorithms↗

Modeling and designing a proteomics application on PROTEUS.

OBJECTIVES: Biomedical applications, such as analysis and management of mass spectrometry proteomics experiments, involve heterogeneous platforms and knowledge, massive data sets, and complex algorithms. Main requirements of such applications are semantic modeling of the experiments and data analysis, as well as high performance computational platforms. In this paper we propose a software platform allowing to model and execute biomedical applications on the Grid. METHODS: Computational Grids offer the required computational power, whereas ontologies and workflow help to face the heterogeneity of biomedical applications. In this paper we propose the use of domain ontologies and workflow techniques for modeling biomedical applications, whereas Grid middleware is responsible for high performance execution. As a case study, the modeling of a proteomics experiment is discussed. RESULTS: The main result is the design and first use of PROTEUS, a Grid-based problem-solving environment for biomedical and bioinformatics applications. CONCLUSION: To manage the complexity of biomedical experiments, ontologies help to model applications and to identify appropriate data and algorithms, workflow techniques allow to combine the elements of such applications in a systematic way. Finally, translation of workflow into execution plans allows the exploitation of the computational power of Grids. Along this direction, in this paper we present PROTEUS discussing a real case study in the proteomics domain.

Algorithms↗