Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Analysis of transcriptional regulation of the small leucine rich proteoglycans.

PURPOSE: Small leucine rich proteoglycans (SLRPs) constitute a family of secreted proteoglycans that are important for collagen fibrillogenesis, cellular growth, differentiation, and migration. Ten of the 13 known members of the SLRP gene family are arranged in tandem clusters on human chromosomes 1, 9, and 12. Their syntenic equivalents are on mouse chromosomes 1, 13, and 10, and rat chromosomes 13, 17, and 7. The purpose of this study was to determine whether there is evidence for control elements, which could regulate the expression of these clusters coordinately. METHODS: Promoters were identified using a comparative genomics approach and Genomatix software tools. For each gene a set of human, mouse, and rat orthologous promoters was extracted from genomic sequences. Transcription factor (TF) binding site analysis combined with a literature search was performed using MatInspector and Genomatix' BiblioSphere. Inspection for the presence of interspecies conserved scaffold/matrix attachment regions (S/MARs) was performed using ElDorado annotation lists. DNAseI hypersensitivity assay, chromatin immunoprecipitation (ChIP), and transient transfection experiments were used to validate the results from bioinformatics analysis. RESULTS: Transcription factor binding site analysis combined with a literature search revealed co-citations between several SLRPs and TFs Runx2 and IRF1, indicating that these TFs have potential roles in transcriptional regulation of the SLRP family members. We therefore inspected all of the SLRP promoter sets for matches to IRF factors and Runx factors. Positionally conserved binding sites for the Runt domain TFs were detected in the proximal promoters of chondroadherin (CHAD) and osteomodulin (OMD) genes. Two significant models (two or more transcription factor binding sites arranged in a defined order and orientation within a defined distance range) were derived from these initial promoter sets, the HOX-Runx (homeodomain-Runt domain), and the ETS-FKHD-STAT (erythroblast transformation specific-forkhead-signal transducers and activators of transcription) models. These models were used to scan the genomic sequences of all 13 SLRP genes. The HOX-Runx model was found within the proximal promoter, exon 1, or intron 1 sequences of 11 of the 13 SLRP genes. The ETS-FKHD-STAT model was found in only 5 of these genes. Transient transfections of MG-63 cells and bovine corneal keratocytes with Runx2 isoforms confirmed the relevance of these TFs to expression of several SLRP genes. Distribution of the HOX-Runx and ETS-FKHD-STAT models within 200 kb of genomic sequence on human chromosome 9 and 500 kb sequence on chromosome 12 also were analyzed. Two regions with 3 HOX-Runx matches within a 1,000 bp window were identified on human chromosome 9; one located between OMD and osteoglycin (OGN)/mimecan genes, and the second located upstream of the putative extracellular matrix protein 2 (ECM2) promoter. The intergenic region between OMD and mimecan was shown to coincide with different patterns of DNAse I hypersensitivity sites in MG-63 and U937 cells. ChiP analysis revealed that this region binds Runx2 in U937 cells (mimecan transcript note detectable), but binds Pitx3 in MG-63 cells (expressing high level of mimecan), thereby demonstrating its functional association with mimecan expression. Upon comparing the predictions of S/MARs on the relevant chromosomal context of human chromosomes 9 and 12 and their rodent equivalents, no convincing evidence was found that the tandemly arranged genes build a chromosomal loop. CONCLUSIONS: Twelve of 13 known SLRP genes have at least one HOX-Runx module match in their promoter, exon 1, intron 1, or intergenic region. Although these genes are located in different clusters on different chromosomes, the common HOX-Runx module could be the basis for co-regulated expression.

Animals↗

Bioinformatics for the 'bench biologist': how to find regulatory regions in genomic DNA.

The combination of bioinformatic and biological approaches constitutes a powerful method for identifying gene regulatory elements. High-quality genome sequences are available in public databases for several vertebrate species. Comparative cross-species sequence analysis of these genomes shows considerable conservation of noncoding sequences in DNA. Biological analyses show that an unexpectedly high number of the conserved sequences correspond to functional cis-regulatory regions that influence gene transcription. Because research biologists are often unfamiliar with the bioinformatic resources at their disposal, this commentary discusses how to integrate biological and bioinformatic methods in the discovery of gene regulatory regions and includes a tutorial on widely available comparative genomics programs.

Animals↗

GeneHuggers: database mining and application connectivity tools for subsequence analyses of the human genome.

UNLABELLED: GeneHuggers is a collection of program modules that enables precise selection of subsequence regions from records of the RefSeq human genome database. Subsequence regions can be selected based on diverse criteria, including feature addresses, annotations from LocusLink and UniGene, and results obtained from analyses with homologous subsequence detection programs. GeneHuggers provides functionality to the UNIX operating system that allows customized bioinformatics program development. AVAILABILITY: GeneHuggers source code is available under the GNU general public license and can be downloaded from ftp://ftp.scripps.edu/pub/genehuggers/gh.tar.gz

Abstracting and Indexing↗

Informatics and quantitative analysis in biological imaging.

Biological imaging is now a quantitative technique for probing cellular structure and dynamics and is increasingly used for cell-based screens. However, the bioinformatics tools required for hypothesis-driven analysis of digital images are still immature. We are developing the Open Microscopy Environment (OME) as an informatics solution for the storage and analysis of optical microscope image data. OME aims to automate image analysis, modeling, and mining of large sets of images and specifies a flexible data model, a relational database, and an XML-encoded file standard that is usable by potentially any software tool. With this design, OME provides a first step toward biological image informatics.

Algorithms↗

Characterization of histone (H1B) oxalate binding protein in experimental urolithiasis and bioinformatics approach to study its oxalate interaction.

The rat kidney H1 oxalate binding protein was isolated and purified. Oxalate binds exclusively with H1B fraction of H1 histone. Oxalate binding activity is inhibited by lysine group modifiers such as 4',4'-diisothiostilbene-2,2-disulfonic acid (DIDS) and pyridoxal phosphate and reduced in presence of ATP and ADP. RNA has no effect on oxalate binding activity of H1B whereas DNA inhibits oxalate binding activity. Equilibrium dialysis method showed that H1B oxalate binding protein has two binding sites for oxalate, one with high affinity, other with low affinity. Histone H1B was modeled in silico using Modeller8v1 software tool since experimental structure is not available. In silico interaction studies predict that histone H1B-oxalate interaction take place through lysine121, lysine139, and leucine68. H1B oxalate binding protein is found to be a promoter of calcium oxalate crystal (CaOx) growth. A 10% increase in the promoting activity is observed in hyperoxaluric rat kidney H1B. Interaction of H1B oxalate binding protein with CaOx crystals favors the formation of intertwined calcium oxalate dehydrate (COD) crystals as studied by light microscopy. Intertwined COD crystals and aggregates of COD crystals were more pronounced in the presence of hyperoxalauric H1B.

Amino Acid Sequence↗

Impact of RNA structure on the prediction of donor and acceptor splice sites.

BACKGROUND: gene identification in genomic DNA sequences by computational methods has become an important task in bioinformatics and computational gene prediction tools are now essential components of every genome sequencing project. Prediction of splice sites is a key step of all gene structural prediction algorithms. RESULTS: we sought the role of mRNA secondary structures and their information contents for five vertebrate and plant splice site datasets. We selected 900-nucleotide sequences centered at each (real or decoy) donor and acceptor sites, and predicted their corresponding RNA structures by Vienna software. Then, based on whether the nucleotide is in a stem or not, the conventional four-letter nucleotide alphabet was translated into an eight-letter alphabet. Zero-, first- and second-order Markov models were selected as the signal detection methods. It is shown that applying the eight-letter alphabet compared to the four-letter alphabet considerably increases the accuracy of both donor and acceptor site predictions in case of higher order Markov models. CONCLUSION: Our results imply that RNA structure contains important data and future gene prediction programs can take advantage of such information.

Algorithms↗

Information services of the European Bioinformatics Institute.

The scope of the EBI is focused on providing better services to the scientific community. Technological advancements in the hardware area provide EBI with means of producing data much faster than before, and with greater accuracy since there is now a better technical ability to produce more exhaustive searches through larger indices. Hand in hand with the technological developments, research and development work is continuing on better indexing systems and more efficient ways of establishing and maintaining the future databases. The existing links of communication between EBI and the user community are exploited to study the needs of the scientific community, to provide better services, and to enhance the quality of databases by interpreting user feedback and updates. A very important goal is to enhance the awareness of the scientific (and, maybe even more, the nonscientific) public of the importance of the modern field of bioinformatics and to introduce special meetings and courses, in which more specific subjects will be studied in depth. Another aspect of this goal is to help in constructing special bioinformatics programs in university faculties. In such programs, in contrast to the existing layout, students will pursue studies in a combined environment that provides basic training in biology and in computation. Currently, one of the main problems in the field is that scientists are either biologists, who are self-educated in the field of computers and programming, or computer scientists without sufficient knowledge of biology. It is hoped that a combined program will provide a high level of education in both fields of interest at the appropriate ratios. Building an efficient and friendly interface between the EBI and the user community is the basis for any future development. This aim is achieved by using the most modern server systems while continuously researching newer and better systems and interfaces. This task can never be complete without involvement of the user community by providing feedback to any of EBI's services. A better bioinformatics community is a necessity for any future development of the biological research aiming at a better society.

Amino Acid Sequence↗

"Plasmo2D": an ancillary proteomic tool to aid identification of proteins from Plasmodium falciparum.

Bioinformatics tools to aid gene and protein sequence analysis have become an integral part of biology in the post-genomic era. Release of the Plasmodium falciparum genome sequence has allowed biologists to define the gene and the predicted protein content as well as their sequences in the parasite. Using pI and molecular weight as characteristics unique to each protein, we have developed a bioinformatics tool to aid identification of proteins from Plasmodium falciparum. The tool makes use of a Virtual 2-DE generated by plotting all of the proteins from the Plasmodium database on a pI versus molecular weight scale. Proteins are identified by comparing the position of migration of desired protein spots from an experimental 2-DE and that on a virtual 2-DE. The procedure has been automated in the form of user-friendly software called "Plasmo2D". The tool can be downloaded from http://144.16.89.25/Plasmo2D.zip.

Animals↗

Identification of autophagy-related genes as potential biomarkers correlated with immune infiltration in bipolar disorder: a bioinformatics analysis.

BACKGROUND: Bipolar disorder (BPD) is a kind of manic and depressive phase alternate episodes of serious mental illness, and it is correlated with well-documented cortical brain abnormalities. Emerging evidence supports that autophagy dysfunction in neuronal system contributes to pathophysiological changes in neurological disease. However, the role of autophagy in bipolar disorder has rarely been elucidated. This study aimed to identify the autophagy-related gene as a potential biomarker Correlated to immune infiltration in BPD. METHODS: The microarray dataset GSE23848 and autophagy-related genes (ARGs) were downloaded. Differentially expressed genes (DEGs) between normal and BPD samples were screened using the R software. Machine learning algorithms were performed to screen the significant candidate biomarker from autophagy-related differentially expressed genes (ARDEGs). The correlation between the screened ARDEGs and infiltrating immune cells was explored through correlation analysis. RESULTS: In this study, the autophagy pathway was abundantly enriched and activated in BPD, as indicated by Pathway enrichment analysis. We identified 16 ARDEGs in BPD compared to the normal group. A signature of 4 ARDEGs (ERN1, ATG3, CTSB, and EIF2AK3) was screened. ROC analysis showed that the above genes have good diagnostic performance. In addition, immune correlation analysis considered that the above four genes significantly correlated with immune cells in BPD. CONCLUSIONS: Autophagy - immune cell axis mediates pathophysiological changes in BPD. Four important ARDEGs are prospective to be potential biomarkers associated with immune infiltration in BPD and helpful for the prediction or diagnosis of BPD.

Bipolar Disorder↗

An overview of Ensembl.

Ensembl (http://www.ensembl.org/) is a bioinformatics project to organize biological information around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of individual genomes, and of the synteny and orthology relationships between them. It is also a framework for integration of any biological data that can be mapped onto features derived from the genomic sequence. Ensembl is available as an interactive Web site, a set of flat files, and as a complete, portable open source software system for handling genomes. All data are provided without restriction, and code is freely available. Ensembl's aims are to continue to "widen" this biological integration to include other model organisms relevant to understanding human biology as they become available; to "deepen" this integration to provide an ever more seamless linkage between equivalent components in different species; and to provide further classification of functional elements in the genome that have been previously elusive.

Computational Biology↗

Analysis of circular genome rearrangement by fusions, fissions and block-interchanges.

BACKGROUND: Analysis of genomes evolving via block-interchange events leads to a combinatorial problem of sorting by block-interchanges, which has been studied recently to evaluate the evolutionary relationship in distance between two biological species since block-interchange can be considered as a generalization of transposition. However, for genomes consisting of multiple chromosomes, their evolutionary history should also include events of chromosome fusions and fissions, where fusion merges two chromosomes into one and fission splits a chromosome into two. RESULTS: In this paper, we study the problem of genome rearrangement between two genomes of circular and multiple chromosomes by considering fusion, fission and block-interchange events altogether. By use of permutation groups in algebra, we propose an O (n2) time algorithm to efficiently compute and obtain a minimum series of fusions, fissions and block-interchanges required to transform one circular multi-chromosomal genome into another, where n is the number of genes shared by the two studied genomes. In addition, we have implemented this algorithm as a web server, called FFBI, and have also applied it to analyzing by gene orders the whole genomes of three human Vibrio pathogens, each with multiple and circular chromosomes, to infer their evolutionary relationships. Consequently, our experimental results coincide well with our previous results obtained using the chromosome-by-chromosome comparisons by landmark orders between any two Vibrio chromosomal sequences as well as using the traditional comparative analysis of 16S rRNA sequences. ConclusionFFBI is a useful tool for the bioinformatics analysis of circular and multiple genome rearrangement by fusions, fissions and block-interchanges.

Algorithms↗

Geometry-based flexible and symmetric protein docking.

We present a set of geometric docking algorithms for rigid, flexible, and cyclic symmetry docking. The algorithms are highly efficient and have demonstrated very good performance in CAPRI Rounds 3-5. The flexible docking algorithm, FlexDock, is unique in its ability to handle any number of hinges in the flexible molecule, without degradation in run-time performance, as compared to rigid docking. The algorithm for reconstruction of cyclically symmetric complexes successfully assembles multimolecular complexes satisfying C(n) symmetry for any n in a matter of minutes on a desktop PC. Most of the algorithms presented here are available at the Tel Aviv University Structural Bioinformatics Web server (http://bioinfo3d.cs.tau.ac.il/).

Algorithms↗

Semiparametric efficient estimation of small genetic effects in large-scale population cohorts.

Population genetics seeks to quantify DNA variant associations with traits or diseases, as well as interactions among variants and with environmental factors. Computing millions of estimates in large cohorts in which small effect sizes and tight confidence intervals are expected, necessitates minimizing model-misspecification bias to increase power and control false discoveries. We present TarGene, a unified statistical workflow for the semi-parametric efficient and double robust estimation of genetic effects including $ k $-point interactions among categorical variables in the presence of confounding and weak population dependence. $ k $-point interactions, or Average Interaction Effects (AIEs), are a direct generalization of the usual average treatment effect (ATE). We estimate genetic effects with cross-validated and/or weighted versions of Targeted Minimum Loss-based Estimators (TMLE) and One-Step Estimators (OSE). The effect of dependence among data units on variance estimates is corrected by using sieve plateau variance estimators based on genetic relatedness across the units. We present extensive realistic simulations to demonstrate power, coverage, and control of type I error. Our motivating application is the targeted estimation of genetic effects on trait, including two-point and higher-order gene-gene and gene-environment interactions, in large-scale genomic databases such as UK Biobank and All of Us. All cross-validated and/or weighted TMLE and OSE for the AIE $ k $-point interaction, as well as ATEs, conditional ATEs and functions thereof, are implemented in the general purpose Julia package TMLE.jl. For high-throughput applications in population genomics, we provide the open-source Nextflow pipeline and software TarGene which integrates seamlessly with modern high-performance and cloud computing platforms.

Humans↗

Poxvirus Bioinformatics Resource Center: a comprehensive Poxviridae informational and analytical resource.

The Poxvirus Bioinformatics Resource Center (PBRC) has been established to provide informational and analytical resources to the scientific community to aid research directed at providing a better understanding of the Poxviridae family of viruses. The PBRC was specifically established as the result of the concern that variola virus, the causative agent of smallpox, as well as related viruses, might be utilized as biological weapons. In addition, the PBRC supports research on poxviruses that might be considered new and emerging infectious agents such as monkeypox virus. The PBRC consists of a relational database and web application that supports the data storage, annotation, analysis and information exchange goals of the project. The current release consists of over 35 complete genomic sequences of various genera, species and strains of viruses from the Poxviridae family. Sequence and annotation information for these viruses has been obtained from sequences publicly available from GenBank as well as sequences not yet deposited in GenBank that have been obtained from ongoing sequencing projects. In addition to sequence data, the PBRC provides comprehensive annotation and curation of virus genes; analytical tools to aid in the understanding of the available sequence data, including tools for the comparative analysis of different virus isolates; and visualization tools to help better display the results of various analyses. The PBRC represents the initial development of what will become a more comprehensive Viral Bioinformatics Resource Center for Biodefense that will be one of the National Institute of Allergy and Infectious Diseases' 'Bioinformatics Resource Centers for Biodefense and Emerging or Re-Emerging Infectious Diseases'. The PBRC website is available at http://www.poxvirus.org.

Computational Biology↗

Methods for identifying and mapping recent segmental and gene duplications in eukaryotic genomes.

The aim of this chapter is to provide instruction for analyzing and mapping recent segmental and gene duplications in eukaryotic genomes. We describe a bioinformatics-based approach utilizing computational tools to manage eukaryotic genome sequences to characterize and understand the evolutionary fates and trajectories of duplicated genes. An introduction to bioinformatics tools and programs such as BLAST, Perl, BioPerl, and the GFF specification provides the necessary background to complete this analysis for any eukaryotic genome of interest.

Animals↗

TAMBIS: transparent access to multiple bioinformatics information sources.

UNLABELLED: TAMBIS (Transparent Access to Multiple Bioinformatics Information Sources) is an application that allows biologists to ask rich and complex questions over a range of bioinformatics resources. It is based on a model of the knowledge of the concepts and their relationships in molecular biology and bioinformatics. AVAILABILITY: TAMBIS is available as an applet from http://img.cs.man.ac.uk/tambis SUPPLEMENTARY: A full manual, tutorial and videos can be found at http://img.cs.man.ac.uk/tambis. CONTACT: tambis@cs.man.ac.uk

Computational Biology↗

Tricross : using dot-plots in sequence-id space to detect uncataloged intergenic features.

MOTIVATION: The process of determining the functional sequence content of an organism is confounded by several factors. Large protein coding sequences are relatively easy to find by statistical methods. Smaller proteins however may escape detection due to their size falling below some arbitrary researcher-defined minimum cutoff, or the inability to precisely define a promoter, or translational start (Delcher et al., Nucleic Acids Res., 27, 4636-4641, 1999). Promoter and regulatory sequences themselves are difficult to define due to a significant amount of allowable sequence variation, as well as a probable lack of any completely accurate whole-organismal gene catalogs to date. Finally, certain genes coding functional RNAs may have insufficient structural or sequence constraints to be detectable by normal sequence structure/pattern searching methods (Eddy and Rivas, Bioinformatics, 16, 583-605, 2000). In those cases where there are multiple closely related organisms that have been sequenced, there is additional information that may be used in the investigation of sequence content-that being the possible conserved nature of functional sequences between the organisms. We present a method for the utilization of this conserved information to detect genes and other potentially functional sequences that may be missed by standard ORF-calling, RNA finding, and pattern matching software. The tricross programs produce a multi-way cross comparison of three sets of sequences, determine which are conserved in all three sets, and produce a graphical (Virtual Reality Modelling Language-VRML; (ISO/IEC 14772-1: 1997, VDC), 1997) representation as well as alignments of all sequence triples found. The software can also be applied to a pair of sequence sets, though the noise in the results increases. RESULTS: Tricross has been used to examine the intergenic-sequence content of the three archaeal Pyrococcus genomes to determine the most highly related sequences remaining between the annotated protein and RNA coding sequences. Set to relatively stringent similarity requirements for the search, tricross found 101 intergenic sequences conserved among the three organisms. Interestingly, 29 of these appear to contain members of a family of small RNA molecules (Kiss-Laszlo et al., EMBO J., 17, 797-807, 1998) only recently discovered in the Archaea (Armbruster, OSU, Diss., 1988; Omer et al., Science, 288, 517-522, 2000; Gaspin et al., J. Mol. Biol., 297, 895-906, 2000). While some of the remaining 72 appear to be individual highly conserved promoter sequences, others have no currently known biological significance. Although originally developed to facilitate the examination of intergenic sequences, none of the tricross logic is inherently specific to intergenic sequences. The software can also be applied to gene sequences, and has been used to produce inter-genomic gene order dot-plots for Haemophilus influenzae (Fleischmann et al., Science, 269, 496-512, 1995) versus H.ducreyi (unpublished data), and Neisseria meningiditis Z2491 (serogroup A) (Parkhill et al., Nature, 404, 502-506, 2000) versus Neisseria meningiditis Z58 (serogroup B) (Tettelin et al., Science, 287, 1809-1815, 2000) versus Neisseria gonorrhoeae (Lewis et al., http://micro-gen.ouhsc.edu/, 2000). AVAILABILITY: The tricross software package is available from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html. CONTACT: ray@biosci.ohio-state.edu; daniels.7@osu.edu; munsonr@pediatrics.ohio-state.edu SUPPLEMENTARY INFORMATION: Additional data from the cross-genomic comparisons examined in the discussion section are linked from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html.

Base Sequence↗

Proteome informatics II: bioinformatics for comparative proteomics.

The present review attempts to cover the most recent initiatives directed towards representing, storing, displaying and processing protein-related data suited to undertake "comparative proteomics" studies. Data interpretation is brought into focus. Efforts invested into analysing and interpreting experimental data increasingly express the need for adding meaning. This trend is perceptible in work dedicated to determining ontologies, modelling interaction networks, etc. In parallel, technical advances in computer science are spurred by the development of the Web and the growing need to channel and understand massive volumes of data. Biology benefits from these advances as an application of choice for many generic solutions. Some examples of bioinformatics solutions are discussed and directions for on-going and future work conclude the review.

Algorithms↗