Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Neighborhood enrichment for the identification of antigen-specific T-cell receptors.

Understanding T-cell receptor (TCR) specificity is not only essential for fundamental research, but could open up novel avenues for diagnostics, cancer immunotherapy, and the targeted treatment of autoimmune diseases. The immune system responds to challenges through groups of T-cells with similar TCR sequences. In recent years, searching for TCRs with an enrichment of similar sequences - neighbors - in a TCR repertoire has become a standard procedure for antigen-specific TCR identification. This study provides a systematic comparison of computational algorithms-ALICE, TCRNET, GLIPH2, and tcrdist3-that leverage neighborhood enrichment for antigen-specific TCR identification. Using published murine datasets from Lymphocytic choriomeningitis virus (LCMV) infection and novel datasets from Sputnik V vaccination and Mycobacterium tuberculosis (Mtb) infection, we evaluated the performance of these algorithms. To facilitate reproducible analysis, we developed TCRgrapher, an R library that integrates these pipelines into a user-friendly framework. TCRgrapher enables efficient identification of antigen-specific TCRs from single repertoire snapshots and supports flexible parameter customization. Our comparative analysis revealed that ALICE and TCRNET consistently outperformed GLIPH2 and tcrdist3 across most datasets, achieving higher area under precision-recall curve. While murine datasets provide valuable insights into algorithm performance, caution is advised when extrapolating these results to other species or different experimental conditions. TCRgrapher is freely available on GitHub (https://github.com/KseniaMIPT/tcrgrapher), offering researchers a robust tool for investigating TCR specificity and advancing immunological studies.

Animals↗

Support vector machine classification and validation of cancer tissue samples using microarray expression data.

MOTIVATION: DNA microarray experiments generating thousands of gene expression measurements, are being used to gather information from tissue and cell samples regarding gene expression differences that will be useful in diagnosing disease. We have developed a new method to analyse this kind of data using support vector machines (SVMs). This analysis consists of both classification of the tissue samples, and an exploration of the data for mis-labeled or questionable tissue results. RESULTS: We demonstrate the method in detail on samples consisting of ovarian cancer tissues, normal ovarian tissues, and other normal tissues. The dataset consists of expression experiment results for 97,802 cDNAs for each tissue. As a result of computational analysis, a tissue sample is discovered and confirmed to be wrongly labeled. Upon correction of this mistake and the removal of an outlier, perfect classification of tissues is achieved, but not with high confidence. We identify and analyse a subset of genes from the ovarian dataset whose expression is highly differentiated between the types of tissues. To show robustness of the SVM method, two previously published datasets from other types of tissues or cells are analysed. The results are comparable to those previously obtained. We show that other machine learning methods also perform comparably to the SVM on many of those datasets. AVAILABILITY: The SVM software is available at http://www.cs. columbia.edu/ approximately bgrundy/svm.

Acute Disease↗

The mutual information: detecting and evaluating dependencies between variables.

MOTIVATION: Clustering co-expressed genes usually requires the definition of 'distance' or 'similarity' between measured datasets, the most common choices being Pearson correlation or Euclidean distance. With the size of available datasets steadily increasing, it has become feasible to consider other, more general, definitions as well. One alternative, based on information theory, is the mutual information, providing a general measure of dependencies between variables. While the use of mutual information in cluster analysis and visualization of large-scale gene expression data has been suggested previously, the earlier studies did not focus on comparing different algorithms to estimate the mutual information from finite data. RESULTS: Here we describe and review several approaches to estimate the mutual information from finite datasets. Our findings show that the algorithms used so far may be quite substantially improved upon. In particular when dealing with small datasets, finite sample effects and other sources of potentially misleading results have to be taken into account.

Algorithms↗

Genetic algorithms applied to multi-class prediction for the analysis of gene expression data.

MOTIVATION: An important challenge in the use of large-scale gene expression data for biological classification occurs when the expression dataset being analyzed involves multiple classes. Key issues that need to be addressed under such circumstances are the efficient selection of good predictive gene groups from datasets that are inherently 'noisy', and the development of new methodologies that can enhance the successful classification of these complex datasets. METHODS: We have applied genetic algorithms (GAs) to the problem of multi-class prediction. A GA-based gene selection scheme is described that automatically determines the members of a predictive gene group, as well as the optimal group size, that maximizes classification success using a maximum likelihood (MLHD) classification method. RESULTS: The GA/MLHD-based approach achieves higher classification accuracies than other published predictive methods on the same multi-class test dataset. It also permits substantial feature reduction in classifier genesets without compromising predictive accuracy. We propose that GA-based algorithms may represent a powerful new tool in the analysis and exploration of complex multi-class gene expression data. AVAILABILITY: Supplementary information, data sets and source codes are available at http://www.omniarray.com/bioinformatics/GA.

Algorithms↗

NPM: latent batch effects correction of omics data by nearest-pair matching.

MOTIVATION: Batch effects (BEs) are a predominant source of noise in omics data and often mask real biological signals. BEs remain common in existing datasets. Current methods for BE correction mostly rely on specific assumptions or complex models, and may not detect and adjust BEs adequately, impacting downstream analysis and discovery power. To address these challenges we developed NPM, a nearest-neighbor matching-based method that adjusts BEs and may outperform other methods in a wide range of datasets. RESULTS: We assessed distinct metrics and graphical readouts, and compared our method to commonly used BE correction methods. NPM demonstrates the ability in correcting for BEs, while preserving biological differences. It may outperform other methods based on multiple metrics. Altogether, NPM proves to be a valuable BE correction approach to maximize discovery in biomedical research, with applicability in clinical research where latent BEs are often dominant. AVAILABILITY AND IMPLEMENTATION: NPM is freely available on GitHub (https://github.com/bigomics/NPM) and on Omics Playground (https://bigomics.ch/omics-playground). Computer codes for analyses are available at (https://github.com/bigomics/NPM). The datasets underlying this article are the following: GSE120099, GSE82177, GSE162760, GSE171343, GSE153380, GSE163214, GSE182440, GSE163857, GSE117970, GSE173078, and GSE10846. All these datasets are publicly available and can be freely accessed on the Gene Expression Omnibus repository.

Humans↗

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis↗

Polaris: Polarization of ancestral and derived polymorphic alleles for inferences of extended haplotype homozygosity in human populations.

SUMMARY: Statistical methods that measure the extent of haplotype homozygosity on chromosomes have been highly informative for identifying episodes of recent selection. For example, the integrated haplotype score (iHS) and the extended haplotype homozygosity (EHH) statistics detect long-range haplotype structure around derived and ancestral alleles indicative of classic and soft selective sweeps, respectively. However, to our knowledge, there are currently no publicly available methods that classify ancestral and derived alleles in genomic datasets for the purpose of quantifying the extent of haplotype homozygosity. Here, we introduce the Polaris package, which polarizes chromosomal variants into ancestral and derived alleles and creates corresponding genetic maps for analysis by selscan and HaploSweep, two versatile haplotype-based programs that perform scans for selection. With the input files generated by Polaris, selscan and/or HaploSweep can produce the appropriate sign (either positive or negative) for outlier iHS statistics, enabling users to distinguish between selection on derived or ancestral alleles. In addition, Polaris can convert the numerical output of these analyses into graphical representations of selective sweeps, increasing the functionality of our software. RESULTS: To demonstrate the utility of our approach, we applied the Polaris package to Chromosome 2 in the European Finnish, Middle Eastern Bedouin, and East African Maasai populations. More specifically, we examined the regulatory sequence in intron 13 of the MCM6 gene associated with lactase persistence (i.e. the ability to digest the lactose sugar present in fresh milk), a region of intense interest to human evolutionary geneticists. Our analyses showed that derived alleles (at known enhancers for lactase expression) sit on an extended haplotype background in the Finnish, Bedouin, and Maasai consistent with a classic selective sweep model as determined by iHS and EHH statistics. Importantly, we were able to immediately identify this target allele under selection based on the information generated by our software. We also explored outlier statistics across Chromosome 2 in two distinct datasets from these populations: (i) one containing polarized alleles generated with Polaris and (ii) the other containing unpolarized alleles in the original phased vcf file. Here, we found an excess of outlier statistics on Chromosome 2 in the unpolarized datasets, raising the possibility that a subset of these "hits" of selection may be unreliable. Overall, Polaris is a versatile package that enables users to efficiently explore, interpret, and report signals of recent selection in genomic datasets. AVAILABILITY AND IMPLEMENTATION: The Polaris package is free and open source on GitHub (https://github.com/alisi1989/Polaris) and DropBox (https://www.dropbox.com/scl/fo/mlxizft5267vem9u62qkn/AAnM0qX923zPzQBlPX8iteM?rlkey=uezrp4t2waffpj0nmo1evr320&e=1&st=jaodccws&dl=0).

Haplotypes↗

Tracing regulatory element networks using epigenetic traits to identify key transcription factors: TENET R/Bioconductor package.

SUMMARY: There is a lack of publicly available bioinformatic tools that can be widely used by researchers to identify transcription factors (TFs) that regulate cell type-specific regulatory elements (REs). To address this, we developed the Tracing regulatory Element Networks using Epigenetic Traits (TENET) R/Bioconductor package. By collecting hundreds of histone mark and open chromatin datasets from a variety of cell lines, primary cells, and tissues, and comparing these features along with matched DNA methylation and gene expression data, TENET identifies TFs and REs linked to a specific cell type. Moreover, we developed methods to interrogate findings using motifs, clinical information, and other genomic and chromatin conformation capture datasets, and applied them to pan-cancer data, highlighting TFs and REs associated with ten different cancer types. TENET enables researchers to better characterize the 3D epigenomes of cell types of interest for future clinical applications. AVAILABILITY AND IMPLEMENTATION: TENET is available at http://bioconductor.org/packages/TENET. Curated functional genomic datasets utilized by TENET are available at http://bioconductor.org/packages/TENET.AnnotationHub. Example datasets are available at http://bioconductor.org/packages/TENET.ExperimentHub.

Transcription Factors↗

ChemGenXplore: an interactive tool for exploring and analysing chemical genomic data.

MOTIVATION: Chemical genomics is a powerful high-throughput approach to systematically link phenotypes to genotypes. However, the vast datasets generated remain challenging to explore due to the lack of integrated, interactive tools for visualization and analysis. Existing workflows often require multiple independent software tools, limiting data accessibility and collaboration. Therefore, we created a user-friendly platform that enables efficient exploration and sharing of chemical genomics data. RESULTS: We developed ChemGenXplore, a web-based Shiny application designed to streamline the visualization and analysis of chemical genomic screens. It offers two primary functionalities: one for exploring pre-implemented datasets and another for analysing user-uploaded datasets. ChemGenXplore enables users to visualize phenotypic profiles, assess gene-gene and condition-condition correlations, perform GO and KEGG enrichment analysis, and generate customizable, interactive heatmaps. To further support collaborative research, ChemGenXplore also facilitates the comparative analysis of chemical genomic and other omics datasets. By consolidating these features into a single interactive and accessible tool, ChemGenXplore facilitates data sharing, enhances reproducibility, and promotes collaboration within the research community. AVAILABILITY AND IMPLEMENTATION: ChemGenXplore is freely accessible as a web application at https://chemgenxplore.kaust.edu.sa/. Source code and documentation, including instructions for local installation, are provided on GitHub (https://github.com/Hudaahmadd/ChemGenXplore). A Docker image is also available on DockerHub (https://hub.docker.com/r/hudaahmad/chemgenxplore) to ensure reproducibility and simplify installation.

Software↗

How negative sampling shapes the performance of transcription factor binding site prediction models.

MOTIVATION: Transcription factors (TFs) are key players in gene regulation and development, where they activate and repress gene expression through DNA binding. Predicting transcription factor binding sites (TFBSs) has long been an active area of research, with many deep learning methods developed to tackle this problem. These models are often trained on TF ChIP-seq data, which is generally seen as only providing positive samples. The choice of datasets and negative sampling techniques is a critical yet often overlooked aspect of this work. RESULTS: In this study, we investigate the impact of different negative sampling techniques on TFBS prediction performance. We create high-quality test datasets based on ChIP-seq and ATAC-seq data, where true negatives can be identified as positions that are accessible but not bound by the TF in question. We then train models using various negative sampling techniques, including genomic sampling, shuffling, dinucleotide shuffling, neighborhood sampling, and cell line specific sampling, simulating cases where matching ATAC-seq data is not available. Our results show that, generally, metrics calculated on training datasets give inflated performance scores. Of the tested techniques, genomic sampling of negatives based on similarity to the positives performed by far the best, although still not reaching the performance of baseline models trained on high-quality datasets. Models trained on dinucleotide shuffled negatives performed poorly, despite being a common practice in the field. Our findings highlight the importance of carefully selecting negative sampling techniques for TFBS prediction, as they can significantly impact model performance and the interpretation of results. AVAILABILITY AND IMPLEMENTATION: The code used in this study is available at https://github.com/NatanTourne/TFBS-negatives (DOI: 10.5281/zenodo.18007567).

Binding Sites↗

CROP: a feature-independent context-aware method for CRISPR-Cas9 frameshift prediction.

MOTIVATION: The CRISPR-Cas9 complex has revolutionized genome-editing technologies. By designing a 20 nt-long guide RNA, a Cas9 nuclease can be guided to cleave almost any genomic target site (followed by NGG). The cleavage induces double-stranded DNA breaks, which are then repaired by cellular pathways. Accurate CRISPR-Cas9 repair-outcome prediction is essential for designing guide RNAs with desired genomic effects, such as gene knockout. A central challenge is quantifying the rate of frameshifts, i.e. repair-outcomes that lead to a change in the local length that is not a multiple of three. Previous methods for frameshift-rate prediction were trained on only a few experimental or cellular contexts, mostly relied on manually defined microhomology features, and were limited by sparse features and class labels. RESULTS: We developed CROP, a feature-independent context-aware repair-outcome prediction method. By aggregating specific repair outcomes as Δlength classes, CROP overcomes class sparsity. We designed CROP to work with variable input sequence lengths and output classes to utilize multiple datasets simultaneously. We benchmarked CROP against state-of-the-art repair-outcome prediction methods over 18 datasets, which we curated and standardized from various studies. Across all datasets, CROP outperformed all competing methods in frameshift-rate prediction. We performed cross-experiment and cross-cellular frameshift-rate predictions to investigate the generalizability of repair mechanisms. Finally, we show that CROP learned microhomology principles from raw sequences without explicit feature engineering, establishing an end-to-end architecture for CRISPR-Cas9 repair-outcome prediction that learns from multiple datasets. AVAILABILITY AND IMPLEMENTATION: CROP is available at https://github.com/OrensteinLab/CROP.

CRISPR-Cas Systems↗

Agentomics: an agentic system that autonomously develops novel state-of-the-art solutions for biomedical machine learning tasks.

MOTIVATION: Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack flexibility, while large language models (LLMs) struggle to consistently deliver reproducible machine learning codebases, and existing LLM Agent-powered solutions lag behind human-engineered ML models. RESULTS: Here, we introduce Agentomics, an autonomous LLM-powered agentic system for end-to-end ML experimentation. Given a biomedical dataset, Agentomics implements various ML modeling strategies, and produces a ready-to-use ML model. Agentomics introduces strict validation checkpoints for standard ML development steps, allowing gradual development on top of working code with defined interfaces and validated artifacts. Further, it offers native support for biomedical foundation models that can be leveraged during experimentation. The generic nature of Agentomics allows the user to create ML solutions for a large variety of datasets and use various LLMs. We evaluate Agentomics across 20 datasets from the domains of Protein Engineering, Drug Discovery, and Regulatory Genomics. When benchmarked against other agentic systems, Agentomics outperformed them in all tested domains. When benchmarked against human expert solutions, Agentomics generated novel state-of-the-art models for 11/20 established benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Agentomics is implemented in Python. Source code and documentation are freely available at: https://github.com/BioGeMT/Agentomics-ML.

Machine Learning↗

WxS-QC-a quality control pipeline for human germline short-variant Whole-Genome and Whole-Exome cohorts for population-scale analyses.

SUMMARY: Whole-exome (WES) and whole-genome (WGS) sequencing are rapidly becoming preferred methods for population-scale analysis of the human genetic landscape. However, there are currently no standardized quality control (QC) pipelines for human WES and WGS datasets. In this paper, we present WxS-QC, a powerful, scalable, and convenient pipeline for the QC of human germline short-variant WGS and WES cohorts for population-scale analyses. Our pipeline is suitable for both rare-variant discovery and common-variant association studies. It is based on deeply refactored gnomAD v3 and v4 quality control pipelines, contains several methods we have developed de novo, and is aligned with current best practices in WGS/WES germline cohort QC. We provide all methods in a single codebase, aligned to work together and controlled via a single YAML config, with automatic export of resulting graphs and summary tables, excellent performance and scalability, and comprehensive documentation. The pipeline can run in any UNIX-like environment and can efficiently process cohorts of up to 200 000 whole-exome samples, with the potential to handle bigger datasets. AVAILABILITY AND IMPLEMENTATION: The pipeline code is written in Python using the Hail library and is freely available under the BSD-3 license here: https://github.com/wtsi-hgi/wxs-qc. The detailed description of the pipeline is available in the pipeline documentation: https://github.com/wtsi-hgi/wxs-qc/blob/main/README.md. We also provide an open dataset with all required metadata, which is available at https://wxs-qc-data.cog.sanger.ac.uk/wxs-qc_public_dataset_v3.tar. An example of test dataset analysis is available in the supplementary materials.

Humans↗

ALPINE: a scalable pipeline for comprehensive classification of gene-editing outcomes from long-read amplicon sequencing.

SUMMARY: CRISPR genome editing has enabled precise genetic modification for gene and cell therapies, but edits often produce heterogeneous on-target outcomes, including homology-directed repair (HDR) knock-ins, DNA repair template integrations, and structural variants. Existing tools are frequently limited to short reads or lack viral vector-specific integration categories needed for therapeutic development. Here, we present ALPINE (Amplicon Long-read Pipeline for INtegration Evaluation), a scalable and reproducible pipeline for classifying and quantifying gene-editing outcomes from long-read amplicon sequencing supporting both PacBio HiFi and Oxford Nanopore platforms. ALPINE classifies reads into 10+ categories, including DNA repair vector integration subtypes, and performs variant calling near the gene-edited site with batch, multi-sample reporting. Uniquely, ALPINE can distinguish between cells treated with multiple DNA repair vectors and identify distinct molecular features, such as inverted terminal repeats (ITRs), enabling comprehensive characterization of complex gene editing outcomes. Dual-target benchmarking on simulated datasets demonstrated high accuracy for transgene integration events. Independent validation on public crosslinked-HDR dataset confirmed ALPINE's integration detection capabilities, and application to edited T cell samples demonstrated comprehensive gene-editing outcome profiling. AVAILABILITY: ALPINE is available under MIT license at https://github.com/Maggi-Chen/ALPINE and https://doi.org/10.5281/zenodo.20272510. All analysis scripts and visualization code used in this manuscript are available at https://github.com/Maggi-Chen/ALPINE-manuscript-analysis. Simulated datasets are deposited at Zenodo (https://doi.org/10.5281/zenodo.20260865). Public dataset PRJNA913199 is available through NCBI SRA.

Gene Editing↗

Combining multiple microarray studies and modeling interstudy variation.

We have established a method for systematic integration of multiple microarray datasets. The method was applied to two different sets of cancer profiling studies. The change of gene expression in cancer was expressed as 'effect size', a standardized index measuring the magnitude of a treatment or covariate effect. The effect sizes were combined to obtain the estimate of the overall mean. The statistical significance was determined by a permutation test extended to multiple datasets. It was shown that the data integration promotes the discovery of small but consistent expression changes with increased sensitivity and reliability. The effect size methods provided the efficient modeling framework for addressing interstudy variation as well. Based on the result of homogeneity tests, a fixed effects model was adopted for one set of datasets that had been created in controlled experimental conditions. By contrast, a random effects model was shown to be appropriate for the other set of datasets that had been published by independent groups. We also developed an alternative modeling procedure based on a Bayesian approach, which would offer flexibility and robustness compared to the classical procedure.

Algorithms↗

A comparison of physical mapping algorithms based on the maximum likelihood model.

MOTIVATION: Physical mapping of chromosomes using the maximum likelihood (ML) model is a problem of high computational complexity entailing both discrete optimization to recover the optimal probe order as well as continuous optimization to recover the optimal inter-probe spacings. In this paper, two versions of the genetic algorithm (GA) are proposed, one with heuristic crossover and deterministic replacement and the other with heuristic crossover and stochastic replacement, for the physical mapping problem under the maximum likelihood model. The genetic algorithms are compared with two other discrete optimization approaches, namely simulated annealing (SA) and large-step Markov chains (LSMC), in terms of solution quality and runtime efficiency. RESULTS: The physical mapping algorithms based on the GA, SA and LSMC have been tested using synthetic datasets and real datasets derived from cosmid libraries of the fungus Neurospora crassa. The GA, especially the version with heuristic crossover and stochastic replacement, is shown to consistently outperform the SA-based and LSMC-based physical mapping algorithms in terms of runtime and final solution quality. Experimental results on real datasets and simulated datasets are presented. Further improvements to the GA in the context of physical mapping under the maximum likelihood model are proposed. AVAILABILITY: The software is available upon request from the first author.

Algorithms↗

Class prediction and discovery using gene microarray and proteomics mass spectroscopy data: curses, caveats, cautions.

MOTIVATION: Two practical realities constrain the analysis of microarray data, mass spectra from proteomics, and biomedical infrared or magnetic resonance spectra. One is the 'curse of dimensionality': the number of features characterizing these data is in the thousands or tens of thousands. The other is the 'curse of dataset sparsity': the number of samples is limited. The consequences of these two curses are far-reaching when such data are used to classify the presence or absence of disease. RESULTS: Using very simple classifiers, we show for several publicly available microarray and proteomics datasets how these curses influence classification outcomes. In particular, even if the sample per feature ratio is increased to the recommended 5-10 by feature extraction/reduction methods, dataset sparsity can render any classification result statistically suspect. In addition, several 'optimal' feature sets are typically identifiable for sparse datasets, all producing perfect classification results, both for the training and independent validation sets. This non-uniqueness leads to interpretational difficulties and casts doubt on the biological relevance of any of these 'optimal' feature sets. We suggest an approach to assess the relative quality of apparently equally good classifiers.

Algorithms↗

Predicting subcellular localization of proteins in a hybridization space.

MOTIVATION: The localization of a protein in a cell is closely correlated with its biological function. With the number of sequences entering into databanks rapidly increasing, the importance of developing a powerful high-throughput tool to determine protein subcellular location has become self-evident. In view of this, the Nearest Neighbour Algorithm was developed for predicting the protein subcellular location using the strategy of hybridizing the information derived from the recent development in gene ontology with that from the functional domain composition as well as the pseudo amino acid composition. RESULTS: As a showcase, the same plant and non-plant protein datasets as investigated by the previous investigators were used for demonstration. The overall success rate of the jackknife test for the plant protein dataset was 86%, and that for the non-plant protein dataset 91.2%. These are the highest success rates achieved so far for the two datasets by following a rigorous cross-validation test procedure, suggesting that such a hybrid approach (particularly by incorporating the knowledge of gene ontology) may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology. AVAILABILITY: The software would be made available on sending a request to the authors.

Algorithms↗