Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

A new bioinformatic approach to detect common 3D sites in protein structures.

An innovative bioinformatic method has been designed and implemented to detect similar three-dimensional (3D) sites in proteins. This approach allows the comparison of protein structures or substructures and detects local spatial similarities: this method is completely independent from the amino acid sequence and from the backbone structure. In contrast to already existing tools, the basis for this method is a representation of the protein structure by a set of stereochemical groups that are defined independently from the notion of amino acid. An efficient heuristic for finding similarities that uses graphs of triangles of chemical groups to represent the protein structures has been developed. The implementation of this heuristic constitutes a software named SuMo (Surfing the Molecules), which allows the dynamic definition of chemical groups, the selection of sites in the proteins, and the management and screening of databases. To show the relevance of this approach, we focused on two extreme examples illustrating convergent and divergent evolution. In two unrelated serine proteases, SuMo detects one common site, which corresponds to the catalytic triad. In the legume lectins family composed of >100 structures that share similar sequences and folds but may have lost their ability to bind a carbohydrate molecule, SuMo discriminates between functional and non-functional lectins with a selectivity of 96%. The time needed for searching a given site in a protein structure is typically 0.1 s on a PIII 800MHz/Linux computer; thus, in further studies, SuMo will be used to screen the PDB.

Algorithms↗

Blast2GO goes grid: developing a grid-enabled prototype for functional genomics analysis.

The vast amount in complexity of data generated in Genomic Research implies that new dedicated and powerful computational tools need to be developed to meet their analysis requirements. Blast2GO (B2G) is a bioinformatics tool for Gene Ontology-based DNA or protein sequence annotation and function-based data mining. The application has been developed with the aim of affering an easy-to-use tool for functional genomics research. Typical B2G users are middle size genomics labs carrying out sequencing, ETS and microarray projects, handling datasets up to several thousand sequences. In the current version of B2G. The power and analytical potential of both annotation and function data-mining is somehow restricted to the computational power behind each particular installation. In order to be able to offer the possibility of an enhanced computational capacity within this bioinformatics application, a Grid component is being developed. A prototype has been conceived for the particular problem of speeding up the Blast searches to obtain fast results for large datasets. Many efforts have been done in the literature concerning the speeding up of Blast searches, but few of them deal with the use of large heterogeneous production Grid Infrastructures. These are the infrastructures that could reach the largest number of resources and the best load balancing for data access. The Grid Service under development will analyse requests based on the number of sequences, splitting them accordingly to the available resources. Lower-level computation will be performed through MPIBLAST. The software architecture is based on the WSRF standard.

Computational Biology↗

MolliGen, a database dedicated to the comparative genomics of Mollicutes.

Bacteria belonging to the class Mollicutes were among the first ones to be selected for complete genome sequencing because of the minimal size of their genomes and their pathogenicity for humans and a broad range of animals and plants. At this time six genome sequences have been publicly released (Mycoplasma genitalium, Mycoplasma pneumoniae, Ureaplasma urealyticum-parvum, Mycoplasma pulmonis, Mycoplasma penetrans and Mycoplasma gallisepticum) and as the number of available mollicute genomes increases, comparative genomics analysis within this model group of organisms becomes more and more instructive. However, such an analysis is difficult to carry out without a suitable platform gathering not only the original annotations but also relevant information available in public databases or obtained by applying common bioinformatics methods. With the aim of solving these difficulties, we have developed a web-accessible database named MolliGen (http://cbi.labri.fr/outils/molligen/). After selecting a set of genomes the user can launch various types of search based on annotation, position on the chromosomes or sequence similarity. In addition, relationships of putative orthology have been precomputed to allow differential genome queries. The results are presented in table format with multiple links to public databases and to bioinformatic analyses such as multiple alignments or BLAST search. Specific tools were also developed for the graphical visualization of the results, including a multi- genome browser for displaying dynamic pictures with clickable objects and for viewing relationships of precomputed similarity. MolliGen is designed to integrate all the complete genomes of mollicutes as they become available.

Computational Biology↗

Proteomics: recent applications and new technologies.

Interest in proteomics as a tool for drug development and a myriad of other applications continues to expand at a rapid rate. Proteomic analyses have recently been conducted on tissues, biofluids, subcellular components and enzymatic pathways as well as various disease and toxicological states, in both animal models and man. In addition, several recent studies have attempted to integrate proteomics data with genomics and/or metabonomics data in a systems biology approach. The translation of proteomic technology and bioinformatics tools to clinical samples, such as in the areas of disease and toxicity biomarkers, represents one of the major opportunities and challenges facing this field. An ongoing challenge in proteomics continues to be the analysis of the serum proteome due to the vast number and complexity of proteins estimated to be present in this biofluid. Aside from the removal of the most abundant proteins, a number of interesting approaches have recently been suggested that may help reduce the overall complexity of serum analysis. In keeping with the increasing interest in applications of proteomics, the tools available for proteomic analyses continue to improve and expand. For example, enhanced tools (such as software and labeling procedures) continue to be developed for the analysis of 2D gels and protein quantification. In addition, activity-based probes are now being used to tag, enrich and isolate distinct sets of proteins based on enzymatic activity. One of the most active areas of development involves microarrays. Antibody-based microarrays have recently been released as commercial products while numerous additional capture agents (e.g. aptamers) and many additional types of microarrays are being explored.

Animals↗

Molecular epidemiological study on pre-X region of hepatitis B virus and identification of hepatocyte proteins interacting with whole-X protein by yeast two-hybrid.

AIM: To identify the pre-X region in hepatitis B virus (HBV) genome and to study the relationship between the genotype and the pre-X region. To investigate the biological function of whole-X (pre-X plus X) protein, we performed yeast two-hybrid to screen proteins in liver interacting with whole-X protein. METHODS: The pre-X region of HBV was amplified by polymerase chain reaction (PCR) method, and was cloned to pGEM Teasy vector. After the target region was sequenced, Vector 8.0 software was used to analyze the sequences. The whole-X bait plasmid was constructed by using yeast two-hybrid system 3. Yeast strain AH109 was transformed. After expression of the whole-X protein in AH109 yeast strains was proved, yeast two-hybrid screening was performed by mating AH109 with Y187 containing liver cDNA library plasmid. The mated yeast was plated on quadruple dropout medium and assayed for alpha-gal activity. The interaction between whole-X protein and the protein obtained from positive colonies was further confirmed by repeating yeast two-hybrid. After extracting and sequencing of plasmid from blue colonies, we carried out analysis by bioinformatics. RESULTS: After sequencing, 27 of 45 clones (60%) were found encoding the pre-X peptide. Eighteen of twenty-seven clones (66.7%) of pre-X coding sequences were found from genotype C. Five positive colonies that interacted with whole-X protein were obtained and sequenced; namely, fetuin B, UDP glycosyltransferase 1 family-polypeptide A9, mannose-P-dolichol utilization defect 1, fibrinogen-B beta polypeptide, transmembrane 4 superfamily member 4-CD81 (TM4SF4). CONCLUSION: The pre-X gene exists in HBV genome. Genes of proteins interacting with whole-X protein in hepatocytes were successfully cloned. These results brought some new clues for studying the biological functions of whole-X protein.

Base Sequence↗

Exploring protein domain structure.

The protein databank contains coordinates of over 10,000 protein structures, which constitute more than 25,000 structural domains in total. The investigation of protein structural, functional and evolutionary relationships is fundamental to many important fields in bioinformatics research, and will be crucial in determining the function of the human and other genomes. This review describes the SCOP and CATH databases of protein structure classification, which define, classify and annotate each domain in the protein databank. The hierarchical structure, use and annotation of the databases are explained. Other tools for exploring protein structure relationships are also described.

Computational Biology↗

Characterization of renal allograft rejection by urinary proteomic analysis.

OBJECTIVE: To develop a diagnostic method with no morbidity or mortality for the detection of acute renal transplant rejection. SUMMARY BACKGROUND DATA: Rejection constitutes the major impediment to the success of transplantation. Currently available methods, including clinical presentation and biochemical organ function parameters, often fail to detect rejection until late stages of progression. Renal biopsies have associated morbidity and mortality and provide only a limited sample of the organ. METHODS: Thirty-four urine samples were collected from 32 renal transplant patients at various stages posttransplantation. Samples were collected from 17 transplant recipients with acute rejection and 15 patients with no rejection. Samples from patients less than 4 days posttransplant were omitted from data analysis due to the presence of excessive inflammatory response proteins. Rejection status was confirmed by kidney biopsy. Specimens were analyzed in triplicate using SELDI mass spectrometry. The obtained spectra were subjected to bioinformatic analysis using ProPeak as well as CART (Classification and Regression Tree) algorithms to identify rejection biomarker candidates. These candidates were identified by their molecular weight and ranked by their ability to distinguish between nonrejection and rejection based on receiver operating characteristic (ROC) analysis. The candidates with the highest area under the ROC curve (AUC) exhibited the best diagnostic performance. RESULTS: The best candidate biomarkers demonstrated highly successful diagnostic performance: 6.5 kd (AUC = 0.839, P <.0001), 6.7 kd (AUC = 0.839, P <.0001), 6.6 kd (AUC = 0.807, P <.0001), 7.1 kd (AUC = 0.807, P <.0001), and 13.4 kd (AUC = 0.804, P <.0001). A separate analysis using the CART algorithm in the Ciphergen Biomarker Pattern Software correctly classified 91% of the 34 specimens in the training set, giving a sensitivity of 83% and specificity of 100% using two separate biomarker candidates at 10.0 kd and 3.4 kd. CONCLUSIONS: Biomarker candidates exist in urine that have the ability to distinguish between renal transplant patients with no rejection and those with acute rejection. These biomarker candidates are the basis for development of a noninvasive method of diagnosing acute rejection without the morbidity and mortality associated with needle biopsy. The combination of biomarkers into a panel for diagnosis leads to the possibility of enhanced diagnostic performance.

Acute Disease↗

Short fuzzy tandem repeats in genomic sequences, identification, and possible role in regulation of gene expression.

MOTIVATION: Genomic sequences are highly redundant and contain many types of repetitive DNA. Fuzzy tandem repeats (FTRs) are of particular interest. They are found in regulatory regions of eukaryotic genes and are reported to interact with transcription factors. However, accurate assessment of FTR occurrences in different genome segments requires specific algorithm for efficient FTR identification and classification. RESULTS: We have obtained formulas for P-values of FTR occurrence and developed an FTR identification algorithm implemented in TandemSWAN software. Using TandemSWAN we compared the structure and the occurrence of FTRs with short period length (up to 24 bp) in coding and non-coding regions including UTRs, heterochromatic, intergenic and enhancer sequences of Drosophila melanogaster and Drosophila pseudoobscura. Tandems with period three and its multiples were found in coding segments, whereas FTRs with periods multiple of six are overrepresented in all non-coding segment. Periods equal to 5-7 and 11-14 were characteristic of the enhancer regions and other non-coding regions close to genes. AVAILABILITY: TandemSWAN web page, stand-alone version and documentation can be found at http://bioinform.genetika.ru/projects/swan/www/ SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

The Cinderella story of metabolic profiling: does metabolomics get to go to the functional genomics ball?

To date most global approaches to functional genomics have centred on genomics, transcriptomics and proteomics. However, since a number of high-profile publications, interest in metabolomics, the global profiling of metabolites in a cell, tissue or organism, has been rapidly increasing. A range of analytical techniques, including 1H NMR spectroscopy, gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), Fourier Transform mass spectrometry (FT-MS), high performance liquid chromatography (HPLC) and electrochemical array (EC-array), are required in order to maximize the number of metabolites that can be identified in a matrix. Applications have included phenotyping of yeast, mice and plants, understanding drug toxicity in pharmaceutical drug safety assessment, monitoring tumour treatment regimes and disease diagnosis in human populations. These successes are likely to be built on as other analytical and bioinformatic approaches are developed to fully exploit the information obtained in metabolic profiles. To assist in this process, databases of metabolomic data will be necessary to allow the passage of information between laboratories. In this prospective review, the capabilities of metabolomics in the field of medicine will be assessed in an attempt to predict the impact this 'Cinderella approach' will have at the 'functional genomic ball'.

Animals↗

Linking experimental results, biological networks and sequence analysis methods using Ontologies and Generalised Data Structures.

The structure of a closely integrated data warehouse is described that is designed to link different types and varying numbers of biological networks, sequence analysis methods and experimental results such as those coming from microarrays. The data schema is inspired by a combination of graph based methods and generalised data structures and makes use of ontologies and meta-data. The core idea is to consider and store biological networks as graphs, and to use generalised data structures (GDS) for the storage of further relevant information. This is possible because many biological networks can be stored as graphs: protein interactions, signal transduction networks, metabolic pathways, gene regulatory networks etc. Nodes in biological graphs represent entities such as promoters, proteins, genes and transcripts whereas the edges of such graphs specify how the nodes are related. The semantics of the nodes and edges are defined using ontologies of node and relation types. Besides generic attributes that most biological entities possess (name, attribute description), further information is stored using generalised data structures. By directly linking to underlying sequences (exons, introns, promoters, amino acid sequences) in a systematic way, close interoperability to sequence analysis methods can be achieved. This approach allows us to store, query and update a wide variety of biological information in a way that is semantically compact without requiring changes at the database schema level when new kinds of biological information is added. We describe how this datawarehouse is being implemented by extending the text-mining framework ONDEX to link, support and complement different bioinformatics applications and research activities such as microarray analysis, sequence analysis and modelling/simulation of biological systems. The system is developed under the GPL license and can be downloaded from http://sourceforge.net/projects/ondex/

Algorithms↗

Non-small cell lung cancer and tumor-educated platelets: screening of biomarkers and construction of a prognostic model.

BACKGROUND: Lung cancer is a leading cause of cancer-related mortality worldwide, emphasizing the urgent need for effective early detection strategies. Traditional Chinese medicine (TCM) provides a unique perspective on tumor pathogenesis, focusing on concepts such as "long-term stasis leading to accumulation". Tumor-educated platelets (TEPs) offer potential as biomarkers due to their ability to reflect cancer heterogeneity and facilitate less invasive diagnostic approaches. This study aims to identify TEP-related prognostic biomarkers for non-small cell lung cancer (NSCLC) and to construct and validate a multigene prognostic model by integrating platelet transcriptomic data with tumor tissue datasets. METHODS: We performed comprehensive analysis of gene expression datasets obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories to characterize transcriptomic differences among lung cancer specimens, normal tissue samples, and TEPs. Using R software, we identified Differentially expressed genes (DEGs) and subsequently applied a multi-stage analytical pipeline to TEP-associated DEGs, incorporating univariate Cox proportional hazards regression, least absolute shrinkage and selection operator (LASSO) regression, multivariate Cox regression, and stepwise regression modeling to pinpoint genes with prognostic significance. These prognostically relevant genes served as the foundation for developing a risk stratification model. We computed individual risk scores across both training and validation cohorts, enabling patient stratification into high- and low-risk categories. Model robustness was assessed through internal cross-validation and external validation procedures, while predictive performance was quantified using risk calibration metrics and receiver operating characteristic (ROC) curve analysis. RESULTS: Through systematic bioinformatics screening, we identified a four-gene prognostic signature comprising NELL2, C4orf48, PRAM1, and KLHL35, which served as the foundation for developing our risk stratification algorithm. Rigorous internal cross-validation and external cohort validation substantiated the moderate predictive performance of this signature. Comprehensive clinicopathological correlation analysis revealed that elevated risk indices, advanced pathological staging (stage III-IV), increased primary tumor dimensions, regional lymph node metastasis, and distant organ dissemination each demonstrated statistically significant associations with diminished overall survival (OS) outcomes in lung cancer patients. The clinical nomogram exhibited acceptable calibration, with calibration plots showing reasonable concordance between predicted and observed survival probabilities across all time points. Discriminative capacity assessment via time-dependent ROC analysis yielded area under the curve (AUC) values consistently surpassing 0.6, confirming moderate prognostic discrimination. Furthermore, decision curve analysis (DCA) demonstrated that our integrated multi-gene model conferred potential net clinical benefit compared to individual prognostic variables across the full spectrum of clinically relevant threshold probabilities (0-1 range), thereby establishing its potential utility for risk-informed clinical decision-making. CONCLUSIONS: This study identified NELL2, C4orf48, PRAM1, and KLHL35 as candidate TEP-related prognostic biomarkers for non-small cell lung cancer (NSCLC). The developed prognostic model shows preliminary potential for patient stratification, but its clinical application, particularly as a platelet-based liquid biopsy tool, requires further validation in independent TEP-based cohorts.

Tumor-educated platelets (TEPs)↗

Selection of target sites for mobile DNA integration in the human genome.

DNA sequences from retroviruses, retrotransposons, DNA transposons, and parvoviruses can all become integrated into the human genome. Accumulation of such sequences accounts for at least 40% of our genome today. These integrating elements are also of interest as gene-delivery vectors for human gene therapy. Here we present a comprehensive bioinformatic analysis of integration targeting by HIV, MLV, ASLV, SFV, L1, SB, and AAV. We used a mathematical method which allowed annotation of each base pair in the human genome for its likelihood of hosting an integration event by each type of element, taking advantage of more than 200 types of genomic annotation. This bioinformatic resource documents a wealth of new associations between genomic features and integration targeting. The study also revealed that the length of genomic intervals analyzed strongly affected the conclusions drawn--thus, answering the question "What genomic features affect integration?" requires carefully specifying the length scale of interest.

Binding Sites↗

Inferential literacy for experimental high-throughput biology.

Many biologists believe that data analysis expertise lags behind the capacity for producing high-throughput data. One view within the bioinformatics community is that biological scientists need to develop algorithmic skills to meet the demands of the new technologies. In this article, we argue that the broader concept of inferential literacy, which includes understanding of data characteristics, experimental design and statistical analysis, in addition to computation, more adequately encompasses what is needed for efficient progress in high-throughput biology.

Animals↗

A serum proteomic pattern for the detection of colorectal adenocarcinoma using surface enhanced laser desorption and ionization mass spectrometry.

PURPOSE: New serum biomarkers are needed to improve the early detection of colorectal adenocarcinoma. We performed surface enhanced laser desorption and ionization time-of-flight mass spectrometry (SELDI-TOF-MS) to screen for differentially expressed proteins in serum and build a proteomic diagnostic pattern for the detection of colorectal adenocarcinoma to improve the prognosis of patients with this disease. EXPERIMENTAL DESIGN: In an attempt to improve current approaches to the serologic diagnosis of colorectal cancer, we analyzed serum samples from subjects with or without colorectal cancer using SELDI-MS. Using a case-control study design, SELDI-MS profile of serum samples from 74 colorectal adenocarcinoma patients were compared with 48 age-and sex-matched healthy subjects using a ProteinChip reader, PBSII-C. Proteomic MS spectra were generated using IMAC3 chips, and protein peaks clustering and classification analyses were performed to build a proteomic pattern that could differentiate patients with colorectal adenocarcinoma from healthy subjects utilizing Biomarker Wizard and Biomarker Patterns software packages, respectively. The constructed pattern was then used to test an independent set of masked serum samples from 60 colorectal cancer patients and 39 healthy subjects. RESULTS: Among the differentially expressed protein peaks identified by SELDI-MS profiling that had the ability to distinguish between patients and healthy subjects, we determined a minimum set of two protein peaks for system training and for developing a decision classification pattern. Masked analysis of an independent set of serum samples showed the diagnostic pattern could differentiate patients with different stages of colorectal cancer from healthy subjects with a sensitivity of 95.00 percent and specificity of 94.87 percent. CONCLUSION: SELDI-TOF-MS profiling of serum proteins combined with bioinformatics tools can be applied to accurately differentiate patients with colorectal cancer from healthy subjects. The high sensitivity and specificity achieved by the constructed clustering analysis algorithm show great potential for the early detection of colorectal cancer.

Adenocarcinoma↗

ProteoMix: an integrated and flexible system for interactively analyzing large numbers of protein sequences.

UNLABELLED: ProteoMix is a suite of JAVA programs for identifying, annotating and predicting regions of interest in large sets of amino acid sequences, according to systematic and consistent criteria. It is based on two concepts (1) the integration of results from different sequence analysis tools increases the prediction reliability; and (2) the integration protocol is critical and needs to be easily adaptable in a case-by-case manner. ProteoMix was designed to analyze simultaneously multiple protein sequences using several bioinformatics tools, merge the results of the analyses using logical functions and display them on an integrated viewer. In addition, new sequences can be added seamlessly to an analysis performed on an initial set of sequences. ProteoMix has a modular design, and bioinformatics tools are run on remote servers accessed using the Internet Simple Object Access Protocol (SOAP), ensuring the swift implementation of additional tools. ProteoMix has a user-friendly interactive graphical user interface environment and runs on PCs with Microsoft OS. AVAILABILITY: ProteoMix is freely available for academic users at http://bio.gsc.riken.jp/ProteoMix/

Database Management Systems↗

The World-Wide Web: an interface between research and teaching in bioinformatics.

The rapid expansion occurring in World-Wide Web activity is beginning to make the concepts of 'global hypermedia' and 'universal document readership realistic objectives of the new revolution in information technology. One consequence of this increase in usage is that educators and students are becoming more aware of the diversity of the knowledge base which can be accessed via the Internet. Although computerised databases and information services have long played a key role in bioinformatics these same resources can also be used to provide core materials for teaching and learning. The large datasets and archives that have been compiled for biomedical research can be enhanced with the addition of a variety of multimedia elements (images, digital videos, animation etc.). The use of this digitally stored information in structured and self-directed learning environments is likely to increase as activity across World-Wide Web increases.

Information Systems↗

An evidence ontology for use in pathway/genome databases.

An important emerging need in Model Organism Databases (MODs) and other bioinformatics databases (DBs) is that of capturing the scientific evidence that supports the information within a DB. This need has become particularly acute as more DB content consists of computationally predicted information, such as predicted gene functions, operons, metabolic pathways, and protein properties. This paper presents an ontology for encoding the type of support and the degree of support for DB assertions, and for encoding the literature source in which that support is reported. The ontology includes a hierarchy of 35 evidence codes for modeling different types of wet-lab and computational evidence for the existence of operons and metabolic pathways, and for gene functions. We also describe an implementation of the ontology within the Pathway Tools software environment, which is used to query and update Pathway/Genome DBs such as EcoCyc, MetaCyc, and HumanCyc.

Computational Biology↗

CRSD: a comprehensive web server for composite regulatory signature discovery.

Transcription factors (TFs) and microRNAs play important roles in the regulation of human gene expression, and the study of their combinatory regulations of gene expression is a new research field. We constructed a comprehensive web server, the composite regulatory signature database (CRSD), that can be applied in investigating complex regulatory behaviors involving gene expression signatures (GESs), microRNA regulatory signatures (MRSs) and TF regulatory signatures (TRSs). Six well-known and large-scale databases, including the human UniGene, mature microRNAs, putative promoter, TRANSFAC, pathway and Gene Ontology (GO) databases, were integrated to provide the comprehensive analysis in CRSD. Two new genome-wide databases, of MRSs and TRSs, were also constructed and further integrated into CRSD. To accomplish the microarray data analysis at one go, several methods, including microarray data pretreatment, statistical and clustering analysis, iterative enrichment analysis and motif discovery, were closely integrated in the web server, which has not been the case in previous studies. Our implementation showed that the published literature could demonstrate the results of genome-wide enrichment analysis. We conclude that CRSD is a powerful and useful bioinformatic web server and may provide new insights into gene regulation networks. CRSD and the online tutorial are publicly available at http://biochip.nchu.edu.tw/crsd1/.

3' Untranslated Regions↗