Search PubMed⌕ Search

Biomedical subjects

Jianjun Hu

Publications and source records attributed to Jianjun Hu.

8 recordsLinked to original sources

Integrative missing value estimation for microarray data.

BACKGROUND: Missing value estimation is an important preprocessing step in microarray analysis. Although several methods have been developed to solve this problem, their performance is unsatisfactory for datasets with high rates of missing data, high measurement noise, or limited numbers of samples. In fact, more than 80% of the time-series datasets in Stanford Microarray Database contain less than eight samples. RESULTS: We present the integrative Missing Value Estimation method (iMISS) by incorporating information from multiple reference microarray datasets to improve missing value estimation. For each gene with missing data, we derive a consistent neighbor-gene list by taking reference data sets into consideration. To determine whether the given reference data sets are sufficiently informative for integration, we use a submatrix imputation approach. Our experiments showed that iMISS can significantly and consistently improve the accuracy of the state-of-the-art Local Least Square (LLS) imputation algorithm by up to 15% improvement in our benchmark tests. CONCLUSION: We demonstrated that the order-statistics-based integrative imputation algorithms can achieve significant improvements over the state-of-the-art missing value estimation approaches such as LLS and is especially good for imputing microarray datasets with a limited number of samples, high rates of missing data, or very noisy measurements. With the rapid accumulation of microarray datasets, the performance of our approach can be further improved by incorporating larger and more appropriate reference datasets.

Algorithms↗

EMD: an ensemble algorithm for discovering regulatory motifs in DNA sequences.

BACKGROUND: Understanding gene regulatory networks has become one of the central research problems in bioinformatics. More than thirty algorithms have been proposed to identify DNA regulatory sites during the past thirty years. However, the prediction accuracy of these algorithms is still quite low. Ensemble algorithms have emerged as an effective strategy in bioinformatics for improving the prediction accuracy by exploiting the synergetic prediction capability of multiple algorithms. RESULTS: We proposed a novel clustering-based ensemble algorithm named EMD for de novo motif discovery by combining multiple predictions from multiple runs of one or more base component algorithms. The ensemble approach is applied to the motif discovery problem for the first time. The algorithm is tested on a benchmark dataset generated from E. coli RegulonDB. The EMD algorithm has achieved 22.4% improvement in terms of the nucleotide level prediction accuracy over the best stand-alone component algorithm. The advantage of the EMD algorithm is more significant for shorter input sequences, but most importantly, it always outperforms or at least stays at the same performance level of the stand-alone component algorithms even for longer sequences. CONCLUSION: We proposed an ensemble approach for the motif discovery problem by taking advantage of the availability of a large number of motif discovery programs. We have shown that the ensemble approach is an effective strategy for improving both sensitivity and specificity, thus the accuracy of the prediction. The advantage of the EMD algorithm is its flexibility in the sense that a new powerful algorithm can be easily added to the system.

Algorithms↗

Integrative Array Analyzer: a software package for analysis of cross-platform and cross-species microarray data.

The rapid accumulation of microarray data translates into an urgent need for tools to perform integrative microarray analysis. Integrative Array Analyzer is a comprehensive analysis and visualization software toolkit, which aims to facilitate the reuse of the large amount of cross-platform and cross-species microarray data. It is composed of the data preprocess module, the co-expression analysis module, the differential expression analysis module, the functional and transcriptional annotation module and the graph visualization module.

Algorithms↗

Limitations and potentials of current motif discovery algorithms.

Computational methods for de novo identification of gene regulation elements, such as transcription factor binding sites, have proved to be useful for deciphering genetic regulatory networks. However, despite the availability of a large number of algorithms, their strengths and weaknesses are not sufficiently understood. Here, we designed a comprehensive set of performance measures and benchmarked five modern sequence-based motif discovery algorithms using large datasets generated from Escherichia coli RegulonDB. Factors that affect the prediction accuracy, scalability and reliability are characterized. It is revealed that the nucleotide and the binding site level accuracy are very low, while the motif level accuracy is relatively high, which indicates that the algorithms can usually capture at least one correct motif in an input sequence. To exploit diverse predictions from multiple runs of one or more algorithms, a consensus ensemble algorithm has been developed, which achieved 6-45% improvement over the base algorithms by increasing both the sensitivity and specificity. Our study illustrates limitations and potentials of existing sequence-based motif discovery algorithms. Taking advantage of the revealed potentials, several promising directions for further improvements are discussed. Since the sequence-based algorithms are the baseline of most of the modern motif discovery algorithms, this paper suggests substantial improvements would be possible for them.

Algorithms↗

Synthesis of carbamate-linked lipids for gene delivery.

Series of lipids 1a-d and 2a,b, with carbamate linkages between hydrocarbon chains and ammonium or tertiary amine head, which were pH sensitive, were synthesized for liposome-mediated gene delivery. The variable length of carbon chains and quaternary ammonium or neutral tertiary amine heads allowed to find the structure-function relationship of how these factors affect cationic lipids on gene delivery performance.

Carbamates↗

Synthesis and characterization of a series of carbamate-linked cationic lipids for gene delivery.

A series of pH-sensitive lipids, la-h [(2,3-bis-alkylcarbamoyloxy-propyl)-trialkylammonium halide] and 2a-d (1-di-methylamino-2,3-bis-alkylcarbamoyloxy-propane), with carbamate linkages between the hydrocarbon chains and an ammonium or tertiary amine head, were synthesized for liposome-mediated gene delivery. The variable length of the carbon chains and the quaternary ammonium or neutral tertiary amine heads allowed us to identify the structure-function relationship showing how these factors would affect cationic lipids in gene delivery performance.

Carbamates↗

NestinnegCD24low/- population from fetal Nestin-EGFP transgenic mice enriches the pancreatic endocrine progenitor cells.

OBJECTIVES: To identify whether Nestin-positve cells or Nestin-negative cells in pancreas enrich potential pancreatic stem/progenitor cells. METHODS: We generated transgenic mice carrying enhanced green fluorescent protein (EGFP) under the control of the nestin second-intronic enhancer and subsequently divided their embryonic pancreatic cells into different subpopulations according to the expression of EGFP and CD24 and characterized these subpopulations by in vitro culture. RESULTS: The EGFP expression correlated well with that of endogenous Nestin. Only the NestinCD24 subpopulation was able to proliferate and generate immature islet-like cell clusters in long-term culture. Immature islet-like cell clusters could be induced to differentiate into insulin-, glucagon-, and somatostatin-positive cells. CONCLUSIONS: Pancreatic endocrine stem/progenitor cells are enriched in the NestinCD24 population of embryonic pancreas.

Animals↗

The hierarchical fair competition (HFC) framework for sustainable evolutionary algorithms.

Many current Evolutionary Algorithms (EAs) suffer from a tendency to converge prematurely or stagnate without progress for complex problems. This may be due to the loss of or failure to discover certain valuable genetic material or the loss of the capability to discover new genetic material before convergence has limited the algorithm's ability to search widely. In this paper, the Hierarchical Fair Competition (HFC) model, including several variants, is proposed as a generic framework for sustainable evolutionary search by transforming the convergent nature of the current EA framework into a non-convergent search process. That is, the structure of HFC does not allow the convergence of the population to the vicinity of any set of optimal or locally optimal solutions. The sustainable search capability of HFC is achieved by ensuring a continuous supply and the incorporation of genetic material in a hierarchical manner, and by culturing and maintaining, but continually renewing, populations of individuals of intermediate fitness levels. HFC employs an assembly-line structure in which subpopulations are hierarchically organized into different fitness levels, reducing the selection pressure within each subpopulation while maintaining the global selection pressure to help ensure the exploitation of the good genetic material found. Three EAs based on the HFC principle are tested - two on the even-10-parity genetic programming benchmark problem and a real-world analog circuit synthesis problem, and another on the HIFF genetic algorithm (GA) benchmark problem. The significant gain in robustness, scalability and efficiency by HFC, with little additional computing effort, and its tolerance of small population sizes, demonstrates its effectiveness on these problems and shows promise of its potential for improving other existing EAs for difficult problems. A paradigm shift from that of most EAs is proposed: rather than trying to escape from local optima or delay convergence at a local optimum, HFC allows the emergence of new optima continually in a bottom-up manner, maintaining low local selection pressure at all fitness levels, while fostering exploitation of high-fitness individuals through promotion to higher levels.

Algorithms↗