Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Bioinformatics: analysis of HLA sequence data.

High resolution HLA typing is a requirement for the selection of matched donors in hematopoietic cell transplantation. The high resolution typing method that provides the most accurate and complete identification of HLA genotypes is sequencing-based typing (SBT). For each sample being tested, SBT defines the exact nucleotide sequence of the coding regions of both alleles at a given HLA locus. Identification of the underlying genotype of the sample can then be made by computerized sequence comparison with all possible HLA allele combinations at that locus. The use of SBT to identify the complete nucleotide sequence of a given HLA gene also enables the direct detection of previously undefined alleles. Since different HLA alleles may differ by a single nucleotide, the accurate assignment of an HLA genotype by SBT is absolutely dependent on the correct identification of the nucleotide at each position for a given sample. However, automated sequence analysis of heterozygous samples may result in the ambiguous assignment of nucleotides at a given position. In addition, ambiguous assignments may result from the sequencing of two different samples that express different HLA alleles but whose sequence profiles appear exactly the same. Both of these ambiguous situations can be resolved by the application of the multi-sequence analysis (MSA) method described here.

Alleles↗

A relational database for management of flow cytometry and ELISpot clinical trial data.

BACKGROUND: Although relational databases are widely used in bioinformatics with deposited and finalized data, they have not received widespread usage among immunologists for managing raw laboratory data such as that generated by ELISpot or flow cytometry assays. Almost no published guidance exists for immunologists to design appropriate and useful data management systems. METHODS: We describe the design and implementation of a Microsoft Access relational database used in a clinical trial in which the primary immunogenicity measures were ELISpot and intracellular cytokine staining. RESULTS: Our data management system enabled us to perform sophisticated queries and to interpret our data as quantitatively as possible. It could easily be used without modification by other researchers using automated plate reading of ELISpot plates or four color flow cytometry. CONCLUSION: We illustrate in detail the use of a flexible data management system for two of the most widely used immunological techniques. Minor modifications for more colors or other outputs can easily be implemented. Based on this example, other modifications could be easily envisaged for any other quantitative output.

Biological Specimen Banks↗

Friend, an integrated analytical front-end application for bioinformatics.

UNLABELLED: Friend is a bioinformatics application designed for simultaneous analysis and visualization of multiple structures and sequences of proteins and/or DNA/RNA. The application provides basic functionalities, such as structure visualization, with different rendering and coloring, sequence alignment and simple phylogeny analysis, along with a number of extended features to perform more complex analyses of sequence structure relationships, including structural alignment of proteins, investigation of specific interaction motifs, studies of protein-protein and protein-DNA interactions and protein super-families. It is also useful for functional annotation of proteins, protein modeling and protein folding studies. Friend provides three levels of usage: (1) an extensive GUI for a scientist with no programming experience, (2) a command line interface for scripting for a scientist with some programming experience and (3) the ability to extend Friend with user written libraries for an experienced programmer. The application is linked and communicates with local and remote sequence and structure databases. AVAILABILITY: http://mozart.bio.neu.edu/friend.

Computational Biology↗

A comparative guide to gene prediction tools for the bioinformatics amateur.

Several hundred programs using different algorithms have been designed to predict individual coding features within any genomic sequence, but none of these tools covers all aspects of a gene or is 100% accurate in its prediction. Automated simultaneous processing of the results from a number of these programs minimizes the chance of a false positive prediction and quickly generates integrated data. We report here on the analysis of two known genes in 5 and 25 kb segments of genomic sequence using four genome annotation packages, NIX, RUMMAGE, Genotator and EMBOSS. Gene predictions were confirmed using cDNA sequences and a comparison was made between the packages. This study showed a similarity in the ability of NIX, RUMMAGE and Genotator to predict well-characterised genes and basic structures, but poor exon prediction for a small, 3 exon gene. However, the BLAST subprograms of all three packages correctly identified the 3 exons. In addition, EST BLAST subprograms identified a previously undescribed, possible 5' untranslated exon for the smaller gene and a number of putative alternatively spliced exons in the larger gene. Overall, NIX was found to be the most user-friendly package, in terms of easy access to databases and the interactive graphical display of results.

Algorithms↗

Prediction of missed cleavage sites in tryptic peptides aids protein identification in proteomics.

Protein identification via peptide mass fingerprinting (PMF) remains a key component of high-throughput proteomics experiments in post-genomic science. Candidate protein identifications are made using bioinformatic tools from peptide peak lists obtained via mass spectrometry (MS). These algorithms rely on several search parameters, including the number of potential uncut peptide bonds matching the primary specificity of the hydrolytic enzyme used in the experiment. Typically, up to one of these "missed cleavages" are considered by the bioinformatics search tools, usually after digestion of the in silico proteome by trypsin. Using two distinct, nonredundant datasets of peptides identified via PMF and tandem MS, a simple predictive method based on information theory is presented which is able to identify experimentally defined missed cleavages with up to 90% accuracy from amino acid sequence alone. Using this simple protocol, we are able to "mask" candidate protein databases so that confident missed cleavage sites need not be considered for in silico digestion. We show that that this leads to an improvement in database searching, with two different search engines, using the PMF dataset as a test set. In addition, the improved approach is also demonstrated on an independent PMF data set of known proteins that also has corresponding high-quality tandem MS data, validating the protein identifications. This approach has wider applicability for proteomics database searching, and the program for predicting missed cleavages and masking Fasta-formatted protein sequence databases has been made available via http:// ispider.smith.man.ac uk/MissedCleave.

Algorithms↗

SPEM: improving multiple sequence alignment with sequence profiles and predicted secondary structures.

MOTIVATION: Multiple sequence alignment is an essential part of bioinformatics tools for a genome-scale study of genes and their evolution relations. However, making an accurate alignment between remote homologs is challenging. Here, we develop a method, called SPEM, that aligns multiple sequences using pre-processed sequence profiles and predicted secondary structures for pairwise alignment, consistency-based scoring for refinement of the pairwise alignment and a progressive algorithm for final multiple alignment. RESULTS: The alignment accuracy of SPEM is compared with those of established methods such as ClustalW, T-Coffee, MUSCLE, ProbCons and PRALINE(PSI) in easy (homologs) and hard (remote homologs) benchmarks. Results indicate that the average sum of pairwise alignment scores given by SPEM are 7-15% higher than those of the methods compared in aligning remote homologs (sequence identity <30%). Its accuracy for aligning homologs (sequence identity >30%) is statistically indistinguishable from those of the state-of-the-art techniques such as ProbCons or MUSCLE 6.0. AVAILABILITY: The SPEM server and its executables are available on http://theory.med.buffalo.edu.

Algorithms↗

Architecture of the herpes simplex virus major capsid protein derived from structural bioinformatics.

The dispositions of 39 alpha helices of greater than 2.5 turns and four beta sheets in the major capsid protein (VP5, 149 kDa) of herpes simplex virus type 1 were identified by computational and visualization analysis from the 8.5A electron cryomicroscopy structure of the whole capsid. The assignment of helices in the VP5 upper domain was validated by comparison with the recently determined crystal structure of this region. Analysis of the spatial arrangement of helices in the middle domain of VP5 revealed that the organization of a tightly associated bundle of ten helices closely resembled that of a domain fold found in the annexin family of proteins. Structure-based sequence searches suggested that sequences in both the N and C-terminal portions of the VP5 sequence contribute to this domain. The long helices seen in the floor domain of VP5 form an interconnected network within and across capsomeres. The combined structural and sequence-based informatics has led to an architectural model of VP5. This model placed in the context of the capsid provides insights into the strategies used to achieve viral capsid stability.

Amino Acid Motifs↗

A map of WW domain family interactions.

WW domains are protein modules that bind proline-rich ligands. WW domain-ligand complexes are of importance as they have been implicated in several human diseases such as muscular dystrophy, cancer, hypertension, Alzheimer's, and Huntington's diseases. We report the results of a protein array aimed at mapping all the human WW domain protein-protein interactions. Our biochemical approach integrates parallel synthesis of peptides, protein expression, and high-throughput screening methodology combined with tools of bioinformatics. The results suggest that the majority of the bioinformatically predicted WW peptide ligands and most WW domains are functional, and that only about 10% of the measured domain-ligand interactions are positive. The analysis of the WW domain protein arrays also underscores the importance of the amino acid residues surrounding the WW ligand core motifs for specific binding to WW domains. In addition, the methodology presented here allows for the rapid elucidation of WW domain-ligand interactions with multiple applications including prediction of exact WW ligand binding sites, which can be applied to the mapping of other protein signaling domain families. Such information can be applied to the generation of protein interaction networks and identification of potential drug targets. To our knowledge, this report describes the first protein-protein interaction map of a domain in the human proteome.

Amino Acid Motifs↗

Postgenomic bioinformatic analysis of yeast artificial chromosome sequence.

The free availability of multiple genomic sequences represents one of the greatest advances in biology of the new millennium, and promises to revolutionize our ability to determine and treat the causes of human disease. This chapter highlights a number of basic, freely available, and user-friendly bioinformatic techniques that can be used to predict the functional genetic contents of specific yeast artificial chromosome (YAC) clones. The content of this chapter is written for the level of graduate students, who may be relatively inexperienced with the use of computers for analyzing DNA sequences. The basic instructions that allow the identification of the genomic sequence of interest and to download this sequence onto a personal computer from an online database are presented. Simple instructions are also given on how to perform basic sequence manipulations, how to use online tools to design polymerase chain reaction primers, and how to map restriction sites. Also described are more complicated programs that rapidly and efficiently perform genome alignments that, in addition to predicting the location of protein coding sequences, allow the prediction of functional genomic sequences, such as cis regulatory elements and scaffold/matrix attachment sites. The availability of genomic sequences and the rapidly expanding numbers of predictive programs that allow the predictive analysis of these sequences promises to greatly facilitate the use of YAC clones in the search for the causes of disease.

Chromosomes, Artificial, Yeast↗

The SH3 domain of nebulin binds selectively to type II peptides: theoretical prediction and experimental validation.

Nebulin, a giant modular protein from muscle, is thought to act as a molecular ruler in sarcomere assembly. The C terminus of nebulin, located in the sarcomere Z-disk, comprises an SH3 domain, a module well known for its role in protein/protein interactions. SH3 domains are known to recognize proline-rich ligands, which have been classified as type I or type II, depending on their relative orientation with respect to the SH3 domain in the complex formed. Type I ligands are bound with their N terminus at the RT loop of the SH3 domain, while type II ligands are bound with their C terminus at the RT loop. Many SH3 domains can bind peptides of either class. Despite the potential importance of the SH3 domain for the function of nebulin as an integral part of a complex network of interactions, no in vivo partner has been identified so far. We have adopted an integrated approach, which combines bioinformatic tools with experimental validation to identify possible partners of nebulin SH3. Using the program SPOT, we performed an exhaustive screening of the muscle sequence databases. This search identified a number of potential nebulin SH3 partners, which were then tested experimentally for their binding affinity. Synthetic peptides were studied by both fluorescence and NMR spectroscopy. Our results show that nebulin SH3 domain binds selectively to type II peptides. The affinity for a type II peptide, 12 residues long, spanning the sequence of a stretch of titin known to colocalise with nebulin in the Z-disk is in the submicromolar range (0.7 microM). This affinity is among the highest found for SH3/peptide complexes, suggesting that the identified stretch could have significance in vivo. The strategy outlined here is of more general applicability and may provide a valuable tool to identify potential partners of SH3 domains and of other peptide-binding modules.

Amino Acid Sequence↗

Tractor workflow: a scalable Nextflow framework for local ancestry-aware genome-wide association studies.

MOTIVATION: The routine exclusion of admixed individuals from traditional genome-wide association studies (GWAS) due to concerns about spurious associations has limited multi-ancestry genetic discovery. Tractor addresses this issue by incorporating local ancestry into association testing, enabling the identification of ancestry-enriched signals and generating ancestry-specific summary statistics. However, adoption has been constrained by the complexity of prerequisite steps, including phasing and local ancestry inference, which require substantial bioinformatics expertise and introduce key analytical decision points. RESULTS: We developed a scalable, automated Nextflow workflow that integrates phasing, local ancestry inference, and Tractor association testing into a reproducible end-to-end pipeline. To demonstrate its utility, we applied the workflow to 32 blood biomarkers in 6245 two-way African-European admixed individuals from the UK Biobank. This pipeline performed efficiently at scale, replicating known associations and uncovering key ancestry-specific loci. These associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously masked genetic signals. AVAILABILITY AND IMPLEMENTATION: The workflow is modular, customizable, and compatible with commonly used phasing and local ancestry tools, minimizing manual intervention while preserving analytical flexibility. By lowering technical barriers to implementation, this framework facilitates broader adoption of local ancestry-aware GWAS, paving the way for expanded genetic discovery.

Humans↗

Identification and characterization of FBXL19 gene in silico.

CXXC1, CXXC2 (FBXL10), CXXC3 (MBD1), CXXC4 (IDAX), CXXC5, CXXC6, CXXC7 (MLL), CXXC8 (FBXL11), CXXC9 (DNMT1) and CXXC10 are CXXC family genes within the human genome. Recently, we identified and characterized CXXC5 and CXXC10 genes as the homologs of CXXC4, which is implicated in the WNT signaling pathway. Here, we identified human FBXL19 (CXXC11) gene by using bioinformatics. Complete coding sequence of FBXL19 cDNA was determined by assembling 10 exons within AC135048.2 genome sequence. NM_019085.1 cDNA was a 5'-truncated partial cDNA corresponding to nucleotide position 138-2025 of FBXL19 complete coding sequence. FBXL19-BCL7C locus at chromosome 16p11.2, FBXL10-RHOF-BCL7A locus at chromosome 12q24.31, and FBXL11-RHOD locus at chromosome 11q13.2 were paralogous regions within the human genome. FBXL19 gene was found to encode a 674-amino-acid FBXL19 protein. Human FBXL19 showed 97.5% total-amino-acid identity with mouse Fbxl19. FBXHA domain (codon 11-128 of FBXL19) and FBXHB domain (codon 404-674 of FBXL19) were identified as novel domains conserved among FBXL19, FBXL10 and FBXL11. CXXC domain was located within the FBXHA domain, and F-box domain was located within the FBXHB domain. FBXL19 consists of FBXHA and FBXHB domains, while FBXL10 and FBXL11 consist of Jumonji C (JmjC), FBXHA and FBXHB domains. This is the first report on human FBXL19 gene as well as FBXHA and FBXHB domains.

Amino Acid Sequence↗

Identification of a novel human glutathione S-transferase using bioinformatics.

In searching the expressed sequence tag (EST) data-base of GenBank with coding sequences of 11 known human glutathione S-transferases in conjunction with bioinformatic analysis, we have identified five ESTs that encode a new human glutathione S-transferase (GST) designated GST A4. The cDNA clone (I.M.A.G.E. Consortium cDNA Clone ID 515157) had an insert length of 1279 bp and contains an open reading frame of 666 bp, which encodes a protein of 222 amino acid residues. The GST A4 protein is identical in length to human GST A1 and A2 and is 54% identical to human GST A1 and A2. Sequence comparison with other human GSTs suggests that it is a new GST belonging to the alpha class GSTs. Northern blot analysis and EST database searches have demonstrated that the GST A4 mRNA is expressed at a high level in brain, placenta, and skeletal muscle and much lower in lung and liver. Analysis of the sequence tagged site (STS) database indicated that the GST A4 gene is located on chromosome 6. This STS represents a previously unidentified transcript further confirming the novelty of the new sequence.

Amino Acid Sequence↗

rSNP_Guide: an integrated database-tools system for studying SNPs and site-directed mutations in transcription factor binding sites.

Since the human genome was sequenced in draft, single nucleotide polymorphism (SNP) analysis has become one of the keynote fields of bioinformatics. We have developed an integrated database-tools system, rSNP_Guide (http://wwwmgs.bionet.nsc.ru/mgs/systems/rsnp/), devoted to prediction of transcription factor (TF) binding sites, alterations of which could be associated with disease phenotype. By inputting data on alterations in DNA sequence and in DNA binding pattern of an unknown TF, rSNP_Guide searches for a known TF with alterations in the recognition score calculated on the basis of TF site's sequence and consistent with the input alterations in DNA binding to the unknown TF. Our system has been tested on many relationships between known TF sites and diseases, as well as on site-directed mutagenesis data. Experimental verification of rSNP_Guide system was made on functionally important SNPs in human TDO2and mouse K-ras genes. Additional examples of analysis are reported involving variants in the human gammaA-globin (HBG1), hsp70(HSPA1A), and Factor IX (F9) gene promoters.

Animals↗

Mastering seeds for genomic size nucleotide BLAST searches.

One of the most common activities in bioinformatics is the search for similar sequences. These searches are usually carried out with the help of programs from the NCBI BLAST family. As the majority of searches are routinely performed with default parameters, a question that should be addressed is how reliable the results obtained using the default parameter values are, i.e. what fraction of potential matches have been retrieved by these searches. Our primary focus is on the initial hit parameter, also known as the seed or word, used by the NCBI BLASTn, MegaBLAST and other similar programs in searches for similar nucleotide sequences. We show that the use of default values for the initial hit parameter can have a big negative impact on the proportion of potentially similar sequences that are retrieved. We also show how the hit probability of different seeds varies with the minimum length and similarity of sequences desired to be retrieved and describe methods that help in determining appropriate seeds. The experimental results described in this paper illustrate situations in which these methods are most applicable and also show the relationship between the various BLAST parameters.

Algorithms↗

Identification and interrogation of highly informative single nucleotide polymorphism sets defined by bacterial multilocus sequence typing databases.

A unified, bioinformatics-driven, single nucleotide polymorphism (SNP)-based approach to microbial genotyping has been developed. Multilocus sequence typing (MLST) databases consist of known variants of standardized housekeeping genes. Normally, seven fragments are defined; a sequence type (ST) consists of the variants of these fragments that are found in a particular isolate. A computer program that can identify highly informative sets of SNPs in entire MLST databases has been constructed. The SNPs either define a particular user-specified ST or provide a high value for Simpson's index of diversity (D), and may thus be generally applicable to that species. SNP sets that are diagnostic for Neisseria meningitidis ST-11 and ST-42, and high-D SNP sets for N. meningitidis and Staphylococcus aureus, were identified and real-time PCR methods to interrogate these SNPs were demonstrated. High-D SNP sets were also identified in other MLST databases. This widely applicable approach allows rapid genetic fingerprinting of infectious agents.

Algorithms↗

Molecular modeling of the membrane targeting of phospholipase C pleckstrin homology domains.

Phospholipases C (PLCs) reversibly associate with membranes to hydrolyze phosphatidylinositol-4, 5-bisphosphate (PI[4,5]P(2)) and comprise four main classes: beta, gamma, delta, and epsilon. Most eukaryotic PLCs contain a single, N-terminal pleckstrin homology (PH) domain, which is thought to play an important role in membrane targeting. The structure of a single PLC PH domain, that from PLCdelta1, has been determined; this PH domain binds PI(4,5)P(2) with high affinity and stereospecificity and has served as a paradigm for PH domain functionality. However, experimental studies demonstrate that PH domains from different PLC classes exhibit diverse modes of membrane interaction, reflecting the dissimilarity in their amino acid sequences. To elucidate the structural basis for their differential membrane-binding specificities, we modeled the three-dimensional structures of all mammalian PLC PH domains by using bioinformatic tools and calculated their biophysical properties by using continuum electrostatic approaches. Our computational analysis accounts for a large body of experimental data, provides predictions for those PH domains with unknown functions, and indicates functional roles for regions other than the canonical lipid-binding site identified in the PLCdelta1-PH structure. In particular, our calculations predict that (1). members from each of the four PLC classes exhibit strikingly different electrostatic profiles than those ordinarily observed for PH domains in general, (2). nonspecific electrostatic interactions contribute to the membrane localization of PLCdelta-, PLCgamma-, and PLCbeta-PH domains, and (3). phosphorylation regulates the interaction of PLCbeta-PH with its effectors through electrostatic repulsion. Our molecular models for PH domains from all of the PLC classes clearly demonstrate how a common structural fold can serve as a scaffold for a wide range of surface features and biophysical properties that support distinctive functional roles.

Amino Acid Sequence↗

ERCnet: Phylogenomic Prediction of Interaction Networks in the Presence of Gene Duplication.

Assigning gene function from genome sequences is a rate-limiting step in molecular biology research. A protein's position within an interaction network can potentially provide insights into its molecular mechanisms. Phylogenetic analysis of evolutionary rate covariation (ERC) in protein sequence has been shown to be effective for large-scale prediction of functional relationships and interactions. However, gene duplication, gene loss, and other sources of phylogenetic incongruence are barriers for analyzing ERC on a genome-wide basis. Here, we developed ERCnet, a bioinformatic program designed to overcome these challenges, facilitating efficient all-versus-all ERC analyses for large protein sequence datasets. We simulated proteome datasets and found that ERCnet achieves combined false positive and negative error rates well below 10% and that our novel "branch-by-branch" length measurements outperforms "root-to-tip" approaches in most cases, offering a valuable new strategy for performing ERC. We also compiled a sample set of 35 angiosperm genomes to test the performance of ERCnet on empirical data, including its sensitivity to user-defined analysis parameters such as input dataset size and branch-length measurement strategy. We investigated the overlap between ERCnet runs with different species samples to understand how species number and composition affect predicted interactions and to identify the protein sets that consistently exhibit ERC across angiosperms. Our systematic exploration of the performance of ERCnet provides a roadmap for design of future ERC analyses to predict functional interactions in a wide array of genomic datasets. ERCnet code is freely available at https://github.com/EvanForsythe/ERCnet.

Gene Duplication↗