Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Soybean genomic survey: BAC-end sequences near RFLP and SSR markers.

We are building a framework physical infrastructure across the soybean genome by using SSR (simple sequence repeat) and RFLP (restriction fragment length polymorphism) markers to identify BACs (bacterial artificial chromosomes) from two soybean BAC libraries. The libraries were prepared from two genotypes, each digested with a different restriction enzyme. The BACs identified by each marker were grouped into contigs. We have obtained BAC- end sequence from BACs within each contig. The sequences were analyzed by the University of Minnesota Center for Computational Genomics and Bioinformatics using BLAST algorithms to search nucleotide and protein databases. The SSR-identified BACs had a higher percentage of significant BLAST hits than did the RFLP-identified BACs. This difference was due to a higher percentage of hits to repetitive-type sequences for the SSR-identified BACs that was offset in part, however, by a somewhat larger proportion of RFLP-identified significant hits with similarity to experimentally defined genes and soybean ESTs (expressed sequence tags). These genes represented a wide range of metabolic functions. In these analyses, only repetitive sequences from SSR-identified contigs appeared to be clustered. The BAC-end sequences also allowed us to identify microsynteny between soybean and the model plants Arabidopsis thaliana and Medicago truncatula. This map-based approach to genome sampling provides a means of assaying soybean genome structure and organization.

Algorithms↗

A Bayesian framework for combining gene predictions.

MOTIVATION: Gene identification and gene discovery in new genomic sequences is one of the most timely computational questions addressed by bioinformatics scientists. This computational research has resulted in several systems that have been used successfully in many whole-genome analysis projects. As the number of such systems grows the need for a rigorous way to combine the predictions becomes more essential. RESULTS: In this paper we provide a Bayesian network framework for combining gene predictions from multiple systems. The framework allows us to treat the problem as combining the advice of multiple experts. Previous work in the area used relatively simple ideas such as majority voting. We introduce, for the first time, the use of hidden input/output Markov models for combining gene predictions. We apply the framework to the analysis of the Adh region in Drosophila that has been carefully studied in the context of gene finding and used as a basis for the GASP competition. The main challenge in combination of gene prediction programs is the fact that the systems are relying on similar features such as cod on usage and as a result the predictions are often correlated. We show that our approach is promising to improve the prediction accuracy and provides a systematic and flexible framework for incorporating multiple sources of evidence into gene prediction systems.

Algorithms↗

Simplifying amino acid alphabets by means of a branch and bound algorithm and substitution matrices.

MOTIVATION: Protein and DNA are generally represented by sequences of letters. In a number of circumstances simplified alphabets (where one or more letters would be represented by the same symbol) have proved their potential utility in several fields of bioinformatics including searching for patterns occurring at an unexpected rate, studying protein folding and finding consensus sequences in multiple alignments. The main issue addressed in this paper is the possibility of finding a general approach that would allow an exhaustive analysis of all the possible simplified alphabets, using substitution matrices like PAM and BLOSUM as a measure for scoring. RESULTS: The computational approach presented in this paper has led to a computer program called AlphaSimp (Alphabet Simplifier) that can perform an exhaustive analysis of the possible simplified amino acid alphabets, using a branch and bound algorithm together with standard or user-defined substitution matrices. The program returns a ranked list of the highest-scoring simplified alphabets. When the extent of the simplification is limited and the simplified alphabets are maintained above ten symbols the program is able to complete the analysis in minutes or even seconds on a personal computer. However, the performance becomes worse, taking up to several hours, for highly simplified alphabets. AVAILABILITY: AlphaSimp and other accessory programs are available at http://bioinformatics.cribi.unipd.it/alphasimp

Algorithms↗

GARSA: genomic analysis resources for sequence annotation.

SUMMARY: Growth of genome data and analysis possibilities have brought new levels of difficulty for scientists to understand, integrate and deal with all this ever-increasing information. In this scenario, GARSA has been conceived aiming to facilitate the tasks of integrating, analyzing and presenting genomic information from several bioinformatics tools and genomic databases, in a flexible way. GARSA is a user-friendly web-based system designed to analyze genomic data in the context of a pipeline. EST and GGS data can be analyzed using the system since it accepts (1) chromatograms, (2) download of sequences from GenBank, (3) Fasta files stored locally or (4) a combination of all three. Quality evaluation of chromatograms, vector removing and clusterization are easily performed as part of the pipeline. A number of local and customizable Blast and CDD analyses can be performed as well as Interpro, complemented with phylogeny analyses. GARSA is being used for the analyses of Trypanosoma vivax (GSS and EST), Trypanosoma rangeli (GSS, EST and ORESTES), Bothrops jararaca (EST), Piaractus mesopotamicus (EST) and Lutzomyia longipalpis (EST). AVAILABILITY: The GARSA system is freely available under GPL license (http://www.biowebdb.org/garsa/). For download requests visit http://www.biowebdb.org/garsa/ or contact Dr Alberto Dávila.

Animals↗

A simple iterative approach to parameter optimization.

Various bioinformatics problems require optimizing several different properties simultaneously. For example, in the protein threading problem, a scoring function combines the values for different parameters of possible sequence-to-structure alignments into a single score to allow for unambiguous optimization. In this context, an essential question is how each property should be weighted. As the native structures are known for some sequences, a partial ordering on optimal alignments to other structures, e.g., derived from structural comparisons, may be used to adjust the weights. To resolve the arising interdependence of weights and computed solutions, we propose a heuristic approach: iterating the computation of solutions (here, threading alignments) given the weights and the estimation of optimal weights of the scoring function given these solutions via systematic calibration methods. For our application (i.e., threading), this iterative approach results in structurally meaningful weights that significantly improve performance on both the training and the test data sets. In addition, the optimized parameters show significant improvements on the recognition rate for a grossly enlarged comprehensive benchmark, a modified recognition protocol as well as modified alignment types (local instead of global and profiles instead of single sequences). These results show the general validity of the optimized weights for the given threading program and the associated scoring contributions.

Algorithms↗

PRIMEGENS: robust and efficient design of gene-specific probes for microarray analysis.

MOTIVATION: DNA microarray is a powerful high-throughput tool for studying gene function and regulatory networks. Due to the problem of potential cross hybridization, using full-length genes for microarray construction is not appropriate in some situations. A bioinformatic tool, PRIMEGENS, has recently been developed for the automatic design of PCR primers using DNA fragments that are specific to individual open reading frames (ORFs). RESULTS: PRIMEGENS first carries out a BLAST search for each target ORF against all other ORFs of the genome to quickly identify possible homologous sequences. Then it performs optimal sequence alignment between the target ORF and each of its homologous ORFs using dynamic programming. PRIMEGENS uses the sequence alignments to select gene- specific fragments, and then feeds the fragments to the Primer3 program to design primer pairs for PCR amplification. PRIMEGENS can be run from the command line on Unix/Linux platforms as a stand-alone package or it can be used from a Web interface. The program runs efficiently, and it takes a few seconds per sequence on a typical workstation. PCR primers specific to individual ORFs from Shewanella oneidensis MR-1 and Deinococcus radiodurans R1 have been designed. The PCR amplification results indicate that this method is very efficient and reliable for designing specific probes for microarray analysis.

Algorithms↗

Profiler: an open web platform for multi-omics analysis.

MOTIVATION: High-throughput multi-omics technologies produce increasingly large and heterogeneous datasets that are difficult to analyze without advanced computational expertise. Existing bioinformatics tools are often fragmented or limited to specific omics types, hindering reproducibility and accessibility. There is a critical need for an integrated, user-friendly, and scalable platform capable of supporting multi-omics analyses across different data modalities. RESULTS: We present Profiler, an open-source, modular platform that unifies data import, quality control, preprocessing, statistical testing, machine and deep learning, biomarker discovery, pathway and drug-target enrichment, and survival modeling within a single reproducible environment. Built in Python with Streamlit, Profiler is available as both a web-based platform deployed on high-performance computing and a desktop version for local execution, enabling flexible usage across computational infrastructures. Profiler supports diverse omics modalities, including proteomics, transcriptomics, lipidomics, and electroencephalogram data. Through applications to glioblastoma proteomic, pancancer, and multi-omics datasets, Profiler reproduced known molecular subtypes, revealed potential therapeutic targets, and generated fully traceable analysis reports within minutes. By integrating advanced analytics behind an intuitive interface, Profiler democratizes multi-omics analysis and provides a robust, scalable foundation for systems biology and precision medicine research. AVAILABILITY AND IMPLEMENTATION: Profiler is open-source and freely available via its web platform (https://prism-profiler.univ-lille.fr) and GitHub (web version: https://github.com/yanisZirem/Profiler_v1_requests_datatests, desktop version: https://github.com/yanisZirem/prism-profiler), and archived on Zenodo (DOI: https://doi.org/10.5281/zenodo.17478158).

Software↗

miBLAST: scalable evaluation of a batch of nucleotide sequence queries with BLAST.

A common task in many modern bioinformatics applications is to match a set of nucleotide query sequences against a large sequence dataset. Existing tools, such as BLAST, are designed to evaluate a single query at a time and can be unacceptably slow when the number of sequences in the query set is large. In this paper, we present a new algorithm, called miBLAST, that evaluates such batch workloads efficiently. At the core, miBLAST employs a q-gram filtering and an index join for efficiently detecting similarity between the query sequences and database sequences. This set-oriented technique, which indexes both the query and the database sets, results in substantial performance improvements over existing methods. Our results show that miBLAST is significantly faster than BLAST in many cases. For example, miBLAST aligned 247 965 oligonucleotide sequences in the Affymetrix probe set against the Human UniGene in 1.26 days, compared with 27.27 days with BLAST (an improvement by a factor of 22). The relative performance of miBLAST increases for larger word sizes; however, it decreases for longer queries. miBLAST employs the familiar BLAST statistical model and output format, guaranteeing the same accuracy as BLAST and facilitating a seamless transition for existing BLAST users.

Algorithms↗

Evolutionary history of the uterine serpins.

A bioinformatics analysis was conducted on the four members of the uterine serpin (US) family of serpins. Evolutionary analysis of the protein sequences and 86 homologous serpins by maximum parsimony and distance methods indicated that the uterine serpins proteins form a clade distinct from other serpins. Ancestral sequences were reconstructed throughout the evolutionary tree by parsimony. These suggested that some branches suffered a high ratio of nonsynonymous to synonymous mutations, suggesting episodes of adaptive evolution within the serpin family. Analysis of the sequences by neutral evolutionary distance methods suggested that the uterine serpins diverged from other serpins prior to the divergence of the mammals from other vertebrates. The porcine uterine serpins are paralogs that diverged from a single common ancestor within the Sus genus after pigs separated from other artiodactyls. The uterine serpins contain several protein kinase C and tyrosine kinase phosphorylation sites. These sites may be important for the lymphocyte-inhibitory activity of OvUS if, like other basic proteins, OvUS can cross the cell membrane of an activated lymphocyte. Internalized OvUS could serve as an alternative target to protein kinases important for the mitogenic response to antigens.

Amino Acid Sequence↗

Towards a consensus on datasets and evaluation metrics for developing B-cell epitope prediction tools.

A B-cell epitope is the three-dimensional structure within an antigen that can be bound to the variable region of an antibody. The prediction of B-cell epitopes is highly desirable for various immunological applications, but has presented a set of unique challenges to the bioinformatics and immunology communities. Improving the accuracy of B-cell epitope prediction methods depends on a community consensus on the data and metrics utilized to develop and evaluate such tools. A workshop, sponsored by the National Institute of Allergy and Infectious Disease (NIAID), was recently held in Washington, DC to discuss the current state of the B-cell epitope prediction field. Many of the currently available tools were surveyed and a set of recommendations was devised to facilitate improvements in the currently existing tools and to expedite future tool development. An underlying theme of the recommendations put forth by the panel is increased collaboration among research groups. By developing common datasets, standardized data formats, and the means with which to consolidate information, we hope to greatly enhance the development of B-cell epitope prediction tools.

Animals↗

PONGO: a web server for multiple predictions of all-alpha transmembrane proteins.

The annotation efforts of the BIOSAPIENS European Network of Excellence have generated several distributed annotation systems (DAS) with the aim of integrating Bioinformatics resources and annotating metazoan genomes (http://www.biosapiens.info). In this context, the PONGO DAS server (http://pongo.biocomp.unibo.it) provides the annotation on predictive basis for the all-alpha membrane proteins in the human genome, not only through DAS queries, but also directly using a simple web interface. In order to produce a more comprehensive analysis of the sequence at hand, this annotation is carried out with four selected and high scoring predictors: TMHMM2.0, MEMSAT, PRODIV and ENSEMBLE1.0. The stored and pre-computed predictions for the human proteins can be searched and displayed in a graphical view. However the web service allows the prediction of the topology of any kind of putative membrane proteins, regardless of the organism and more importantly with the same sequence profile for a given sequence when required. Here we present a new web server that incorporates the state-of-the-art topology predictors in a single framework, so that putative users can interactively compare and evaluate four predictions simultaneously for a given sequence. Together with the predicted topology, the server also displays a signal peptide prediction determined with SPEP. The PONGO web server is available at http://pongo.biocomp.unibo.it/pongo.

Humans↗

Human proton/oligopeptide transporter (POT) genes: identification of putative human genes using bioinformatics.

The proton-dependent oligopeptide transporters (POT) gene family currently consists of approximately 70 cloned cDNAs derived from diverse organisms. In mammals, two genes encoding peptide transporters, PepT1 and PepT2 have been cloned in several species including humans, in addition to a rat histidine/peptide transporter (rPHT1). Because the Candida elegans genome contains five putative POT genes, we searched the available protein and nucleic acid databases for additional mammalian/human POT genes, using iterative BLAST runs and the human expressed sequence tags (EST) database. The apparent human orthologue of rPHT1 (expression largely confined to rat brain and retina) was represented by numerous ESTs originating from many tissues. Assembly of these ESTs resulted in a contiguous sequence covering approximately 95% of the suspected coding region. The contig sequences and analyses revealed the presence of several possible splice variants of hPHT1. A second closely related human EST-contig displayed high identity to a recently cloned mouse cDNA encoding cyclic adenosine monophosphate (cAMP)-inducible 1 protein (gi:4580995). This contig served to identify a PAC clone containing deduced exons and introns of the likely human orthologue (termed hPHT2). Northern analyses with EST clones indicated that hPHT1 is primarily expressed in skeletal muscle and spleen, whereas hPHT2 is found in spleen, placenta, lung, leukocytes, and heart. These results suggest considerable complexity of the human POT gene family, with relevance to the absorption and distribution of cephalosporins and other peptoid drugs.

Blotting, Northern↗

Identification and characterization of ASXL3 gene in silico.

Polycomb group proteins are implicated in embryogenesis and carcinogenesis through transcriptional regulation of target genes. ASXL1 and ASXL2 genes, encoding Polycomb group protein with ASXN and ASXM domains, are human homologs of Drosophila additional sex combs (asx) gene. Exons 2-13 of the ASXL2 gene are fused to exons 1-14 of the MYST3 gene in a case of therapy-related myelodysplastic syndrome due to t(2;8)(p23.3;p11.2). Here, we identified the ASXL3 gene, a novel human homolog of Drosophila asx, by using bioinformatics. ASXL3 gene, consisting of 12 exons, was located within human genome sequences RP11-562H1 (AC023192.8), RP11-265C19 (AC090989.8), and RP11-470B24 (AC010798.9). Complete coding sequence of human ASXL3 cDNA was determined by assembling EST BE145544, exons 4-11, and 5'-truncated KIAA1713 cDNA (AB051500.2). Partial coding sequence of mouse Asxl3 cDNA was derived from 3'-truncated C230079D11 cDNA (AK082659.1). Human ASXL3 mRNA was expressed in pancreatic islet, testis as well as in neuroblastoma, head and neck tumor. Human ASXL3 protein (2248 aa) with ASXN, ASXM and PHD domains was the third member of the human ASXL family. The region between ASXM and PHD domains was divergent among ASXL family members. Proline-rich domain was located within the divergent region of ASXL3, but not within that of ASXL1 and ASXL2. ASXL3-DTNA locus at chromosome 18q12.1 and ASXL2-DTNB locus at 2p23.3 were paralogous regions within the human genome. ASXL3 was a predicted cancer-associated gene, just like ASXL1 and ASXL2. This is the first report on identification and characterization of the ASXL3 gene.

Amino Acid Sequence↗

A structure-based method for protein sequence alignment.

MOTIVATION: With the continuing rapid growth of protein sequence data, protein sequence comparison methods have become the most widely used tools of bioinformatics. Among these methods are those that use position-specific scoring matrices (PSSMs) to describe protein families. PSSMs can capture information about conserved patterns within families, which can be used to increase the sensitivity of searches for related sequences. Certain types of structural information, however, are not generally captured by PSSM search methods. Here we introduce a program, Structure-based ALignment TOol (SALTO), that aligns protein query sequences to PSSMs using rules for placing and scoring gaps that are consistent with the conserved regions of domain alignments from NCBI's Conserved Domain Database. RESULTS: In most cases, the alignment scores obtained using the local alignment version follow an extreme value distribution. SALTO's performance in finding related sequences and producing accurate alignments is similar to or better than that of IMPALA; one advantage of SALTO is that it imposes an explicit gapping model on each protein family. AVAILABILITY: A stand-alone version of the program that can generate global or local alignments is available by ftp distribution (ftp://ftp.ncbi.nih.gov/pub/SALTO/), and has been incorporated to Cn3D structure/alignment viewer. CONTACT: bryant@ncbi.nlm.nih.gov.

Algorithms↗

Gene2Oligo: oligonucleotide design for in vitro gene synthesis.

There is substantial interest in implementing a bioinformatics tool that allows the design of oligonucleotides to support the development of in vitro gene synthesis. Current protocols to make long synthetic DNA molecules rely on the in vitro assembly of a set of short oligonucleotides, either by ligase chain reaction (LCR) or by assembly PCR. Ideally, such oligonucleotides should represent both strands of the final DNA molecule. They should be adjacent on the same strand and overlap the complementary oligonucleotides from the second strand to ensure good hybridization during assembly. This implies that the thermodynamic properties of each oligonucleotide have to be consistent across the set. Furthermore, any given oligonucleotide has to be totally specific to its target to avoid the creation of incorrectly assembled sequences. We have developed Gene2Oligo (http://berry.engin.umich.edu/gene2oligo/), a web-based tool that divides a long input DNA sequence into a set of adjacent oligonucleotides representing both DNA strands. The length of the oligonucleotides is dynamically optimized to ensure both the specificity and the uniform melting temperatures necessary for in vitro gene synthesis. We have successfully designed and used a set of oligonucleotides to synthesize the Saccharomyces cerevisiae cytochrome b5 by using both LCR and assembly PCR.

Algorithms↗

Large-scale production of SAGE libraries from microdissected tissues, flow-sorted cells, and cell lines.

We describe the details of a serial analysis of gene expression (SAGE) library construction and analysis platform that has enabled the generation of >298 high-quality SAGE libraries and >30 million SAGE tags primarily from sub-microgram amounts of total RNA purified from samples acquired by microdissection. Several RNA isolation methods were used to handle the diversity of samples processed, and various measures were applied to minimize ditag PCR carryover contamination. Modifications in the SAGE protocol resulted in improved cloning and DNA sequencing efficiencies. Bioinformatic measures to automatically assess DNA sequencing results were implemented to analyze the integrity of ditag structure, linker or cross-species ditag contamination, and yield of high-quality tags per sequence read. Our analysis of singleton tag errors resulted in a method for correcting such errors to statistically determine tag accuracy. From the libraries generated, we produced an essentially complete mapping of reliable 21-base-pair tags to the mouse reference genome sequence for a meta-library of approximately 5 million tags. Our analyses led us to reject the commonly held notion that duplicate ditags are artifacts. Rather than the usual practice of discarding such tags, we conclude that they should be retained to avoid introducing bias into the results and thereby maintain the quantitative nature of the data, which is a major theoretical advantage of SAGE as a tool for global transcriptional profiling.

Animals↗

Bio++: a set of C++ libraries for sequence analysis, phylogenetics, molecular evolution and population genetics.

BACKGROUND: A large number of bioinformatics applications in the fields of bio-sequence analysis, molecular evolution and population genetics typically share input/output methods, data storage requirements and data analysis algorithms. Such common features may be conveniently bundled into re-usable libraries, which enable the rapid development of new methods and robust applications. RESULTS: We present Bio++, a set of Object Oriented libraries written in C++. Available components include classes for data storage and handling (nucleotide/amino-acid/codon sequences, trees, distance matrices, population genetics datasets), various input/output formats, basic sequence manipulation (concatenation, transcription, translation, etc.), phylogenetic analysis (maximum parsimony, markov models, distance methods, likelihood computation and maximization), population genetics/genomics (diversity statistics, neutrality tests, various multi-locus analyses) and various algorithms for numerical calculus. CONCLUSION: Implementation of methods aims at being both efficient and user-friendly. A special concern was given to the library design to enable easy extension and new methods development. We defined a general hierarchy of classes that allow the developer to implement its own algorithms while remaining compatible with the rest of the libraries. Bio++ source code is distributed free of charge under the CeCILL general public licence from its website http://kimura.univ-montp2.fr/BioPP.

Algorithms↗

Columba: fast approximate pattern matching with optimized search schemes.

MOTIVATION: Aligning sequencing reads to reference genomes is a fundamental task in bioinformatics. Aligners can be classified as lossy or lossless: lossy aligners prioritize speed by reporting only one or a few high-scoring alignments, whereas lossless aligners output all optimal alignments, ensuring completeness and sensitivity. RESULTS: This paper introduces Columba, a high-performance lossless aligner tailored for Illumina sequencing data. Columba processes single or paired-end reads in FASTQ format and outputs alignments in SAM format. By utilizing advanced search schemes and bit-parallel alignment techniques, Columba achieves exceptional speed. Columba is available in two variants. The first, based on the bidirectional FM-index, prioritizes speed. The second, Columba RLC, uses run-length compression using a bidirectional move structure, significantly reducing memory usage for large, repetitive datasets like pan-genomes. Benchmarks on the human genome, as well as bacterial and human pan-genome datasets, demonstrate that Columba is much faster than existing lossless aligners and even competitive with lossy tools. We integrated Columba into the OptiType HLA genotyping pipeline, where it substantially reduced computational time while maintaining accuracy. These results position Columba as a versatile, state-of-the-art tool for high-sensitivity genomic analyses. AVAILABILITY AND IMPLEMENTATION: The source code of Columba is available at https://github.com/biointec/columba under AGPL license. Scripts to reproduce the benchmarks and analyses are available at https://doi.org/10.5281/zenodo.15849246.

Software↗