Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

PhyloNaP: a user-friendly database of phylogeny for natural product-producing enzymes.

SUMMARY: Phylogenetic analysis is widely used to predict enzyme function, yet building annotated and reusable trees is labor-intensive and requires extensive knowledge about the specific enzymes. Existing resources rarely cover biosynthetic enzymes and lack the context needed for meaningful analysis. We present PhyloNaP, the first large-scale resource dedicated to phylogenies of biosynthetic enzymes. PhyloNaP provides ∼51 000 annotated and interactive trees enriched with chemical, functional, and taxonomic information. Users can classify their own sequences via phylogenetic placement, enabling functional inference in an evolutionary context. A contribution portal allows the community to submit curated trees. By combining scale, breadth of annotation, and interactive functionality, PhyloNaP fills a major gap in bioinformatics resources for enzyme discovery and annotation, with immediate applications to secondary metabolism and beyond. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at https://phylonap.cs.uni-tuebingen.de.

Phylogeny↗

Interpretation of mass spectrometry data for high-throughput proteomics.

Recent developments in proteomics have revealed a bottleneck in bioinformatics: high-quality interpretation of acquired MS data. The ability to generate thousands of MS spectra per day, and the demand for this, makes manual methods inadequate for analysis and underlines the need to transfer the advanced capabilities of an expert human user into sophisticated MS interpretation algorithms. The identification rate in current high-throughput proteomics studies is not only a matter of instrumentation. We present software for high-throughput PMF identification, which enables robust and confident protein identification at higher rates. This has been achieved by automated calibration, peak rejection, and use of a meta search approach which employs various PMF search engines. The automatic calibration consists of a dynamic, spectral information-dependent algorithm, which combines various known calibration methods and iteratively establishes an optimised calibration. The peak rejection algorithm filters signals that are unrelated to the analysed protein by use of automatically generated and dataset-dependent exclusion lists. In the "meta search" several known PMF search engines are triggered and their results are merged by use of a meta score. The significance of the meta score was assessed by simulation of PMF identification with 10,000 artificial spectra resembling a data situation close to the measured dataset. By means of this simulation the meta score is linked to expectation values as a statistical measure. The presented software is part of the proteome database ProteinScape which links the information derived from MS data to other relevant proteomics data. We demonstrate the performance of the presented system with MS data from 1891 PMF spectra. As a result of automatic calibration and peak rejection the identification rate increased from 6% to 44%.

Algorithms↗

Comprehensive circRNA expression profile and hub genes screening during human liver development.

BACKGROUND: Understanding the expression of non-coding RNA in the liver during embryonic development provides important insights into liver diseases. Therefore, we investigated circular RNA (circRNA) roles in human liver development, an unexplored research domain. METHODS: Using high-throughput sequencing and bioinformatics, we analysed foetal liver samples across developmental stages (7-20 weeks post-conception). Differentially expressed (DE) genes were identified and subjected to enrichment analysis using Gene Ontology (GO), Kyoto Encyclopaedia of Genes and Genomes (KEGG), and Disease Ontology (DO). Modular analysis was performed using the Search Tool for Retrieval of Interacting Genes (STRING), followed by construction of a protein-protein interaction (PPI) network using Cytoscape software. The key genes were screened using Molecular Complex Detection (MCODE). The mRNA levels of hub genes were validated using quantitative reverse transcription polymerase chain reaction (qRT-PCR). RESULTS: There were 645 DE circRNAs and 5,145 DE mRNAs between human livers at the three growth stages (HB, EH, and LH). It was found that the activity of circRNAs was boosted remarkably in the hepatoblastic stage. Enrichment analysis found they mainly involved in nervous system regulation of liver function, embryonic organ development and digestive system development. In addition, DE circRNAs were primarily involved in the PI3K-AKT, MAPK and calcium pathways, potentially contributing to adult liver diseases. Notably, only hsa_circ_001471 and novel_circ_017382 were simultaneously identified at all stages and were persistently downregulated. A co-expression regulatory network involving these circRNAs was established. Three hub genes (LGR5, FOXL1 and RSPO3) were identified from the PPI network of 167 genes and may play key roles in human liver development. The RT-qPCR validation results were in agreement with the sequencing data. CONCLUSIONS: Our findings provide the first insights into the roles and regulatory networks of circRNAs in human liver development, laying the groundwork for further investigations of molecular and signalling networks.

Humans↗

SeqVISTA: a graphical tool for sequence feature visualization and comparison.

BACKGROUND: Many readers will sympathize with the following story. You are viewing a gene sequence in Entrez, and you want to find whether it contains a particular sequence motif. You reach for the browser's "find in page" button, but those darn spaces every 10 bp get in the way. And what if the motif is on the opposite strand? Subsequently, your favorite sequence analysis software informs you that there is an interesting feature at position 13982-14013. By painstakingly counting the 10 bp blocks, you are able to examine the sequence at this location. But now you want to see what other features have been annotated close by, and this information is buried several screenfuls higher up the web page. RESULTS: SeqVISTA presents a holistic, graphical view of features annotated on nucleotide or protein sequences. This interactive tool highlights the residues in the sequence that correspond to features chosen by the user, and allows easy searching for sequence motifs or extraction of particular subsequences. SeqVISTA is able to display results from diverse sequence analysis tools in an integrated fashion, and aims to provide much-needed unity to the bioinformatics resources scattered around the Internet. Our viewer may be launched on a GenBank record by a single click of a button installed in the web browser. CONCLUSION: SeqVISTA allows insights to be gained by viewing the totality of sequence annotations and predictions, which may be more revealing than the sum of their parts. SeqVISTA runs on any operating system with a Java 1.4 virtual machine. It is freely available to academic users at http://zlab.bu.edu/SeqVISTA.

Amino Acid Sequence↗

Informatics and quantitative analysis in biological imaging.

Biological imaging is now a quantitative technique for probing cellular structure and dynamics and is increasingly used for cell-based screens. However, the bioinformatics tools required for hypothesis-driven analysis of digital images are still immature. We are developing the Open Microscopy Environment (OME) as an informatics solution for the storage and analysis of optical microscope image data. OME aims to automate image analysis, modeling, and mining of large sets of images and specifies a flexible data model, a relational database, and an XML-encoded file standard that is usable by potentially any software tool. With this design, OME provides a first step toward biological image informatics.

Algorithms↗

Information services of the European Bioinformatics Institute.

The scope of the EBI is focused on providing better services to the scientific community. Technological advancements in the hardware area provide EBI with means of producing data much faster than before, and with greater accuracy since there is now a better technical ability to produce more exhaustive searches through larger indices. Hand in hand with the technological developments, research and development work is continuing on better indexing systems and more efficient ways of establishing and maintaining the future databases. The existing links of communication between EBI and the user community are exploited to study the needs of the scientific community, to provide better services, and to enhance the quality of databases by interpreting user feedback and updates. A very important goal is to enhance the awareness of the scientific (and, maybe even more, the nonscientific) public of the importance of the modern field of bioinformatics and to introduce special meetings and courses, in which more specific subjects will be studied in depth. Another aspect of this goal is to help in constructing special bioinformatics programs in university faculties. In such programs, in contrast to the existing layout, students will pursue studies in a combined environment that provides basic training in biology and in computation. Currently, one of the main problems in the field is that scientists are either biologists, who are self-educated in the field of computers and programming, or computer scientists without sufficient knowledge of biology. It is hoped that a combined program will provide a high level of education in both fields of interest at the appropriate ratios. Building an efficient and friendly interface between the EBI and the user community is the basis for any future development. This aim is achieved by using the most modern server systems while continuously researching newer and better systems and interfaces. This task can never be complete without involvement of the user community by providing feedback to any of EBI's services. A better bioinformatics community is a necessity for any future development of the biological research aiming at a better society.

Amino Acid Sequence↗

Identification of autophagy-related genes as potential biomarkers correlated with immune infiltration in bipolar disorder: a bioinformatics analysis.

BACKGROUND: Bipolar disorder (BPD) is a kind of manic and depressive phase alternate episodes of serious mental illness, and it is correlated with well-documented cortical brain abnormalities. Emerging evidence supports that autophagy dysfunction in neuronal system contributes to pathophysiological changes in neurological disease. However, the role of autophagy in bipolar disorder has rarely been elucidated. This study aimed to identify the autophagy-related gene as a potential biomarker Correlated to immune infiltration in BPD. METHODS: The microarray dataset GSE23848 and autophagy-related genes (ARGs) were downloaded. Differentially expressed genes (DEGs) between normal and BPD samples were screened using the R software. Machine learning algorithms were performed to screen the significant candidate biomarker from autophagy-related differentially expressed genes (ARDEGs). The correlation between the screened ARDEGs and infiltrating immune cells was explored through correlation analysis. RESULTS: In this study, the autophagy pathway was abundantly enriched and activated in BPD, as indicated by Pathway enrichment analysis. We identified 16 ARDEGs in BPD compared to the normal group. A signature of 4 ARDEGs (ERN1, ATG3, CTSB, and EIF2AK3) was screened. ROC analysis showed that the above genes have good diagnostic performance. In addition, immune correlation analysis considered that the above four genes significantly correlated with immune cells in BPD. CONCLUSIONS: Autophagy - immune cell axis mediates pathophysiological changes in BPD. Four important ARDEGs are prospective to be potential biomarkers associated with immune infiltration in BPD and helpful for the prediction or diagnosis of BPD.

Bipolar Disorder↗

Semiparametric efficient estimation of small genetic effects in large-scale population cohorts.

Population genetics seeks to quantify DNA variant associations with traits or diseases, as well as interactions among variants and with environmental factors. Computing millions of estimates in large cohorts in which small effect sizes and tight confidence intervals are expected, necessitates minimizing model-misspecification bias to increase power and control false discoveries. We present TarGene, a unified statistical workflow for the semi-parametric efficient and double robust estimation of genetic effects including $ k $-point interactions among categorical variables in the presence of confounding and weak population dependence. $ k $-point interactions, or Average Interaction Effects (AIEs), are a direct generalization of the usual average treatment effect (ATE). We estimate genetic effects with cross-validated and/or weighted versions of Targeted Minimum Loss-based Estimators (TMLE) and One-Step Estimators (OSE). The effect of dependence among data units on variance estimates is corrected by using sieve plateau variance estimators based on genetic relatedness across the units. We present extensive realistic simulations to demonstrate power, coverage, and control of type I error. Our motivating application is the targeted estimation of genetic effects on trait, including two-point and higher-order gene-gene and gene-environment interactions, in large-scale genomic databases such as UK Biobank and All of Us. All cross-validated and/or weighted TMLE and OSE for the AIE $ k $-point interaction, as well as ATEs, conditional ATEs and functions thereof, are implemented in the general purpose Julia package TMLE.jl. For high-throughput applications in population genomics, we provide the open-source Nextflow pipeline and software TarGene which integrates seamlessly with modern high-performance and cloud computing platforms.

Humans↗

TAMBIS: transparent access to multiple bioinformatics information sources.

UNLABELLED: TAMBIS (Transparent Access to Multiple Bioinformatics Information Sources) is an application that allows biologists to ask rich and complex questions over a range of bioinformatics resources. It is based on a model of the knowledge of the concepts and their relationships in molecular biology and bioinformatics. AVAILABILITY: TAMBIS is available as an applet from http://img.cs.man.ac.uk/tambis SUPPLEMENTARY: A full manual, tutorial and videos can be found at http://img.cs.man.ac.uk/tambis. CONTACT: tambis@cs.man.ac.uk

Computational Biology↗

Tricross : using dot-plots in sequence-id space to detect uncataloged intergenic features.

MOTIVATION: The process of determining the functional sequence content of an organism is confounded by several factors. Large protein coding sequences are relatively easy to find by statistical methods. Smaller proteins however may escape detection due to their size falling below some arbitrary researcher-defined minimum cutoff, or the inability to precisely define a promoter, or translational start (Delcher et al., Nucleic Acids Res., 27, 4636-4641, 1999). Promoter and regulatory sequences themselves are difficult to define due to a significant amount of allowable sequence variation, as well as a probable lack of any completely accurate whole-organismal gene catalogs to date. Finally, certain genes coding functional RNAs may have insufficient structural or sequence constraints to be detectable by normal sequence structure/pattern searching methods (Eddy and Rivas, Bioinformatics, 16, 583-605, 2000). In those cases where there are multiple closely related organisms that have been sequenced, there is additional information that may be used in the investigation of sequence content-that being the possible conserved nature of functional sequences between the organisms. We present a method for the utilization of this conserved information to detect genes and other potentially functional sequences that may be missed by standard ORF-calling, RNA finding, and pattern matching software. The tricross programs produce a multi-way cross comparison of three sets of sequences, determine which are conserved in all three sets, and produce a graphical (Virtual Reality Modelling Language-VRML; (ISO/IEC 14772-1: 1997, VDC), 1997) representation as well as alignments of all sequence triples found. The software can also be applied to a pair of sequence sets, though the noise in the results increases. RESULTS: Tricross has been used to examine the intergenic-sequence content of the three archaeal Pyrococcus genomes to determine the most highly related sequences remaining between the annotated protein and RNA coding sequences. Set to relatively stringent similarity requirements for the search, tricross found 101 intergenic sequences conserved among the three organisms. Interestingly, 29 of these appear to contain members of a family of small RNA molecules (Kiss-Laszlo et al., EMBO J., 17, 797-807, 1998) only recently discovered in the Archaea (Armbruster, OSU, Diss., 1988; Omer et al., Science, 288, 517-522, 2000; Gaspin et al., J. Mol. Biol., 297, 895-906, 2000). While some of the remaining 72 appear to be individual highly conserved promoter sequences, others have no currently known biological significance. Although originally developed to facilitate the examination of intergenic sequences, none of the tricross logic is inherently specific to intergenic sequences. The software can also be applied to gene sequences, and has been used to produce inter-genomic gene order dot-plots for Haemophilus influenzae (Fleischmann et al., Science, 269, 496-512, 1995) versus H.ducreyi (unpublished data), and Neisseria meningiditis Z2491 (serogroup A) (Parkhill et al., Nature, 404, 502-506, 2000) versus Neisseria meningiditis Z58 (serogroup B) (Tettelin et al., Science, 287, 1809-1815, 2000) versus Neisseria gonorrhoeae (Lewis et al., http://micro-gen.ouhsc.edu/, 2000). AVAILABILITY: The tricross software package is available from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html. CONTACT: ray@biosci.ohio-state.edu; daniels.7@osu.edu; munsonr@pediatrics.ohio-state.edu SUPPLEMENTARY INFORMATION: Additional data from the cross-genomic comparisons examined in the discussion section are linked from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html.

Base Sequence↗

Dehydron: a structurally encoded signal for protein interaction.

We introduce a quantifiable structural motif, called dehydron, that is shown to be central to protein-protein interactions. A dehydron is a defectively packed backbone hydrogen bond suggesting preformed monomeric structure whose Coulomb energy is highly sensitive to binding-induced water exclusion. Such preformed hydrogen bonds are effectively adhesive, since water removal from their vicinity contributes to their stability. At the structural level, a significant correlation is established between dehydrons and sites for protein complexation, with the HIV-1 capsid protein P24 complexed with antibody light-chain FAB25.3 providing the most dramatic correlation. Furthermore, the number of dehydrons in homologous similar-fold proteins from different species is shown to be a signature of proteomic complexity. The techniques are then applied to higher levels of organization: The formation of the capsid and its organization in picornaviruses correlates strongly with the distribution of dehydrons on the rim of the virus unit. Furthermore, antibody contacts and crystal contacts may be assigned to dehydrons still prevalent after the capsid has been assembled. The implications of the dehydron as an encoded signal in proteomics, bioinformatics, and inhibitor drug design are emphasized.

Amino Acid Motifs↗

Bioinformatics in glycobiology.

In comparison with genes and proteins, attention paid to oligosaccharides that modify proteins is still marginal. Accordingly, bioinformatics is so far poorly involved in glycobiology. Some initiatives have been taken, however, to collect in databases all glycobiology-relevant information or to design specific data mining algorithms to infer predictions or identify oligosaccharide structures. In this review, we make a non-exhaustive survey of the available glycobiology-related bioinformatic resources, focussing mainly on those resources that are available through the World Wide Web. Some well-curated databases are identified, but the development of specialised algorithms appears to be limited.

Algorithms↗

UK CropNet: a collection of databases and bioinformatics resources for crop plant genomics.

The UK Crop Plant Bioinformatics Network (UK CropNet) was established in 1996 in order to harness the extensive work in genome mapping in crop plants in the UK. Since this date we have published five databases from our central UK CropNet WWW site (http://synteny.nott.ac.uk/) with a further three to follow shortly. Our resource facilitates the identification and manipulation of agronomically important genes by laying a foundation for comparative analysis among crop plants and model species. In addition, we have developed a number of software tools that facilitate the visualisation and analysis of our data. Many of our tools are made freely available for use with both crop plant data and with data from other species.

Crops, Agricultural↗

Mapping the proteome of Leishmania Viannia parasites using two-dimensional polyacrylamide gel electrophoresis and associated technologies.

In this study we have demonstrated the potential of two-dimensional electrophoresis (2DE)-based technologies as tools for characterization of the Leishmania proteome (the expressed protein complement of the genome). Standardized neutral range (pH 5-7) proteome maps of Leishmania (Viannia) guyanensis and Leishmania (Viannia) panamensis promastigotes were reproducibly generated by 2DE of soluble parasite extracts, which were prepared using lysis buffer containing urea and nonidet P-40 detergent. The Coomassie blue and silver nitrate staining systems both yielded good resolution and representation of protein spots, enabling the detection of approximately 800 and 1,500 distinct proteins, respectively. Several reference protein spots common to the proteomes of all parasite species/strains studied were isolated and identified by peptide mass spectrometry (LC-ES-MS/MS), and bioinformatics approaches as members of the heat shock protein family, ribosomal protein S12, kinetoplast membrane protein 11 and a hypothetical Leishmania-specific 13 kDa protein of unknown function. Immunoblotting of Leishmania protein maps using a monoclonal antibody resulted in the specific detection of the 81.4 kDa and 77.5 kDa subunits of paraflagellar rod proteins 1 and 2, respectively. Moreover, differences in protein expression profiles between distinct parasite clones were reproducibly detected through comparative proteome analyses of paired maps using image analysis software. These data illustrate the resolving power of 2DE-based proteome analysis. The production and basic characterization of good quality Leishmania proteome maps provides an essential first step towards comparative protein expression studies aimed at identifying the molecular determinants of parasite drug resistance and virulence, as well as discovering new drug and vaccine targets.

Animals↗

GoMiner: a resource for biological interpretation of genomic and proteomic data.

We have developed GoMiner, a program package that organizes lists of 'interesting' genes (for example, under- and overexpressed genes from a microarray experiment) for biological interpretation in the context of the Gene Ontology. GoMiner provides quantitative and statistical output files and two useful visualizations. The first is a tree-like structure analogous to that in the AmiGO browser and the second is a compact, dynamically interactive 'directed acyclic graph'. Genes displayed in GoMiner are linked to major public bioinformatics resources.

Computer Graphics↗

Elucidating the Mechanism of Xiaoqinglong Decoction in Chronic Urticaria Treatment: An Integrated Approach of Network Pharmacology, Bioinformatics Analysis, Molecular Docking, and Molecular Dynamics Simulations.

INTRODUCTION: Xiaoqinglong Decoction (XQLD) is a traditional Chinese medicinal formula commonly used to treat chronic urticaria (CU). However, its underlying therapeutic mechanisms remain incompletely characterized. This study employed an integrated approach combining network pharmacology, bioinformatics, molecular docking, and molecular dynamics simulations to identify the active components, potential targets, and related signaling pathways involved in XQLD's therapeutic action against CU, thereby providing a mechanistic foundation for its clinical application. METHODS: The active components of XQLD and their corresponding targets were identified using the Traditional Chinese Medicine Systems Pharmacology (TCMSP) database. CU-related targets were retrieved from the OMIM and GeneCards databases. Subsequently, core components and targets were determined via protein-protein interaction (PPI) network analysis and component-target-pathway network construction. Topological analyses were performed using Cytoscape software to prioritize core nodes within these networks. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were conducted via the DAVID database to identify enriched biological processes and signaling pathways. Molecular docking was performed to evaluate binding interactions between key components and core targets, while molecular dynamics (MD) simulations were employed to assess the stability of the component-target complexes with the lowest binding energy. Finally, CU-related targets of XQLD were validated using datasets from the Gene Expression Omnibus (GEO) database. RESULTS: A total of 135 active components and 249 potential targets of XQLD were identified, alongside 1,711 CU-related targets. Core components, such as quercetin, kaempferol, beta-sitosterol, naringenin, stigmasterol, and luteolin, exhibited high degree values in the constructed networks. The core targets identified included AKT1, TNF, IL6, TP53, PTGS2, CASP3, BCL2, ESR1, PPARG, and MAPK3. GO and KEGG pathway enrichment analyses revealed the PI3K-Akt signaling pathway as a central regulatory mechanism. Molecular docking studies demonstrated strong binding affinities between active components and core targets, with the stigmasterol-AKT1 complex exhibiting the lowest binding energy (-11.4 kcal/mol) and high stability in MD simulations. Validation using GEO datasets identified 12 core genes shared between CU-related targets and XQLD-associated targets, including PTGS2 and IL6, which were also prioritized as core targets in the network pharmacology analyses. DISCUSSION: This study comprehensively integrates multidisciplinary approaches to clarify the potential molecular mechanisms of XQLD in treating CU, highlighting its multitarget and multipathway synergistic effects. Molecular docking and dynamics simulations confirm the stable interaction between stigmasterol and the core target AKT1. Additionally, GEO dataset analysis verifies the pathogenic relevance of targets such as PTGS2 and IL6, significantly enhancing the credibility of our findings. These results provide a modern scientific basis for the traditional therapeutic effects of XQLD on CU and have important implications for developing multitarget treatments for this condition. However, this study mainly relies on database mining and computational simulations. Further in vitro and in vivo experimental validations are needed to confirm the predicted component-target-pathway interactions. CONCLUSION: This study identifies the active components, potential targets, and pathways through which XQLD exerts therapeutic effects on CU. These findings provide a theoretical foundation for further mechanistic studies and support their clinical application in the treatment of CU.

Molecular Docking Simulation↗

A novel splice-altering TNC variant (c.5247A > T, p.Gly1749Gly) in an Chinese family with autosomal dominant non-syndromic hearing loss.

BACKGROUND: This study aims to analyze the pathogenic gene in a Chinese family with non-syndromic hearing loss and identify a novel mutation site in the TNC gene. METHODS: A five-generation Chinese family from Anhui Province, presenting with autosomal dominant non-syndromic hearing loss, was recruited for this study. By analyzing the family history, conducting clinical examinations, and performing genetic analysis, we have thoroughly investigated potential pathogenic factors in this family. The peripheral blood samples were obtained from 20 family members, and the pathogenic genes were identified through whole exome sequencing. Subsequently, the mutation of gene locus was confirmed using Sanger sequencing. The conservation of TNC mutation sites was assessed using Clustal Omega software. We utilized functional prediction software including dbscSNV_AdaBoost, dbscSNV_RandomForest, NNSplice, NetGene2, and Mutation Taster to accurately predict the pathogenicity of these mutations. Furthermore, exon deletions were validated through RT-PCR analysis. RESULTS: The family exhibited autosomal dominant, progressive, post-lingual, non-syndromic hearing loss. A novel synonymous variant (c.5247A > T, p.Gly1749Gly) in TNC was identified in affected members. This variant is situated at the exon-intron junction boundary towards the end of exon 18. Notably, glycine residue at position 1749 is highly conserved across various species. Bioinformatics analysis indicates that this synonymous mutation leads to the disruption of the 5' end donor splicing site in the 18th intron of the TNC gene. Meanwhile, verification experiments have demonstrated that this synonymous mutation disrupts the splicing process of exon 18, leading to complete exon 18 skipping and direct splicing between exons 17 and 19. CONCLUSION: This novel splice-altering variant (c.5247A > T, p.Gly1749Gly) in exon 18 of the TNC gene disrupts normal gene splicing and causes hearing loss among HBD families.

Adult↗