Search PubMed⌕ Search

Biomedical subjects

Shaoqi Rao

Publications and source records attributed to Shaoqi Rao.

16 recordsLinked to original sources

Effects of replacing the unreliable cDNA microarray measurements on the disease classification based on gene expression profiles and functional modules.

MOTIVATION: Microarrays datasets frequently contain a large number of missing values (MVs), which need to be estimated and replaced for subsequent data mining. The focus of the paper is to study the effects of different MV treatments for cDNA microarray data on disease classification analysis. RESULTS: By analyzing five datasets, we demonstrate that among three kinds of classifiers evaluated in this study, support vector machine (SVM) classifiers are robust to varied MV imputation methods [e.g. replacing MVs by zero, K nearest-neighbor (KNN) imputation algorithm, local least square imputation and Bayesian principal component analysis], while the classification and regression tree classifiers are sensitive in terms of classification accuracy. The KNNclassifiers built on differentially expressed genes (DEGs) are robust to the varied MV treatments, but the performances of the KNN classifiers based on all measured genes can be significantly deteriorated when imputing MVs for genes with larger missing rate (MR) (e.g. MR > 5%). Generally, while replacing MVs by zero performs relatively poor, the other imputation algorithms have little difference in affecting classification performances of the SVM or KNN classifiers. We further demonstrate the power and feasibility of our recently proposed functional expression profile (FEP) approach as means to handle microarray data with MVs. The FEPs, which are derived from the functional modules that are enriched with sets of DEGs and thus can be consistently identified under varied MV treatments, achieve precise disease classification with better biological interpretation. We conclude that the choice of MV treatments should be determined in context of the later approaches used for disease classification. The suggested exclusion criterion of ignoring the genes with larger MR (e.g. >5%), while justifiable for some classifiers such as KNN classifiers, might not be considered as a general rule for all classifiers.

Algorithms↗

Discovery of Time-Delayed Gene Regulatory Networks based on temporal gene expression profiling.

BACKGROUND: It is one of the ultimate goals for modern biological research to fully elucidate the intricate interplays and the regulations of the molecular determinants that propel and characterize the progression of versatile life phenomena, to name a few, cell cycling, developmental biology, aging, and the progressive and recurrent pathogenesis of complex diseases. The vast amount of large-scale and genome-wide time-resolved data is becoming increasing available, which provides the golden opportunity to unravel the challenging reverse-engineering problem of time-delayed gene regulatory networks. RESULTS: In particular, this methodological paper aims to reconstruct regulatory networks from temporal gene expression data by using delayed correlations between genes, i.e., pairwise overlaps of expression levels shifted in time relative each other. We have thus developed a novel model-free computational toolbox termed TdGRN (Time-delayed Gene Regulatory Network) to address the underlying regulations of genes that can span any unit(s) of time intervals. This bioinformatics toolbox has provided a unified approach to uncovering time trends of gene regulations through decision analysis of the newly designed time-delayed gene expression matrix. We have applied the proposed method to yeast cell cycling and human HeLa cell cycling and have discovered most of the underlying time-delayed regulations that are supported by multiple lines of experimental evidence and that are remarkably consistent with the current knowledge on phase characteristics for the cell cyclings. CONCLUSION: We established a usable and powerful model-free approach to dissecting high-order dynamic trends of gene-gene interactions. We have carefully validated the proposed algorithm by applying it to two publicly available cell cycling datasets. In addition to uncovering the time trends of gene regulations for cell cycling, this unified approach can also be used to study the complex gene regulations related to the development, aging and progressive pathogenesis of a complex disease where potential dependences between different experiment units might occurs.

Algorithms↗

A novel model-free approach for reconstruction of time-delayed gene regulatory networks.

Reconstruction of genetic networks is one of the key scientific challenges in functional genomics. This paper describes a novel approach for addressing the regulatory dependencies between genes whose activities can be delayed by multiple units of time. The aim of the proposed approach termed TdGRN (time-delayed gene regulatory networking) is to reversely engineer the dynamic mechanisms of gene regulations, which is realized by identifying the time-delayed gene regulations through supervised decision-tree analysis of the newly designed time-delayed gene expression matrix, derived from the original time-series microarray data. A permutation technique is used to determine the statistical classification threshold of a tree, from which a gene regulatory rule(s) is extracted. The proposed TdGRN is a model-free approach that attempts to learn the underlying regulatory rules without relying on any model assumptions. Compared with model-based approaches, it has several significant advantages: it requires neither any arbitrary threshold for discretization of gene transcriptional values nor the definition of the number of regulators (k). We have applied this novel method to the publicly available data for budding yeast cell cycling. The numerical results demonstrate that most of the identified time-delayed gene regulations have current biological knowledge supports.

Algorithms↗

Towards precise classification of cancers based on robust gene functional expression profiles.

BACKGROUND: Development of robust and efficient methods for analyzing and interpreting high dimension gene expression profiles continues to be a focus in computational biology. The accumulated experiment evidence supports the assumption that genes express and perform their functions in modular fashions in cells. Therefore, there is an open space for development of the timely and relevant computational algorithms that use robust functional expression profiles towards precise classification of complex human diseases at the modular level. RESULTS: Inspired by the insight that genes act as a module to carry out a highly integrated cellular function, we thus define a low dimension functional expression profile for data reduction. After annotating each individual gene to functional categories defined in a proper gene function classification system such as Gene Ontology applied in this study, we identify those functional categories enriched with differentially expressed genes. For each functional category or functional module, we compute a summary measure (s) for the raw expression values of the annotated genes to capture the overall activity level of the module. In this way, we can treat the gene expressions within a functional module as an integrative data point to replace the multiple values of individual genes. We compare the classification performance of decision trees based on functional expression profiles with the conventional gene expression profiles using four publicly available datasets, which indicates that precise classification of tumour types and improved interpretation can be achieved with the reduced functional expression profiles. CONCLUSION: This modular approach is demonstrated to be a powerful alternative approach to analyzing high dimension microarray data and is robust to high measurement noise and intrinsic biological variance inherent in microarray data. Furthermore, efficient integration with current biological knowledge has facilitated the interpretation of the underlying molecular mechanisms for complex human diseases at the modular level.

Algorithms↗

A robust hybrid between genetic algorithm and support vector machine for extracting an optimal feature gene subset.

Development of a robust and efficient approach for extracting useful information from microarray data continues to be a significant and challenging task. Microarray data are characterized by a high dimension, high signal-to-noise ratio, and high correlations between genes, but with a relatively small sample size. Current methods for dimensional reduction can further be improved for the scenario of the presence of a single (or a few) high influential gene(s) in which its effect in the feature subset would prohibit inclusion of other important genes. We have formalized a robust gene selection approach based on a hybrid between genetic algorithm and support vector machine. The major goal of this hybridization was to exploit fully their respective merits (e.g., robustness to the size of solution space and capability of handling a very large dimension of feature genes) for identification of key feature genes (or molecular signatures) for a complex biological phenotype. We have applied the approach to the microarray data of diffuse large B cell lymphoma to demonstrate its behaviors and properties for mining the high-dimension data of genome-wide gene expression profiles. The resulting classifier(s) (the optimal gene subset(s)) has achieved the highest accuracy (99%) for prediction of independent microarray samples in comparisons with marginal filters and a hybrid between genetic algorithm and K nearest neighbors.

Algorithms↗

Genome-wide linkage scan identifies a novel genetic locus on chromosome 5p13 for neonatal atrial fibrillation associated with sudden death and variable cardiomyopathy.

BACKGROUND: Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia, and patients with AF have a significantly increased risk for ischemic stroke. Approximately 15% of all strokes are caused by AF. The molecular basis and underlying mechanisms and pathophysiology of AF remain largely unknown. METHODS AND RESULTS: We have identified a large AF family with an autosomal recessive inheritance pattern. The AF in the family manifests with early onset at the fetal stage and is associated with neonatal sudden death and, in some cases, ventricular tachyarrhythmias and waxing and waning cardiomyopathy. Genome-wide linkage analysis was performed for 36 family members and generated a 2-point logarithm of the odds (LOD) score of 3.05 for marker D5S455. The maximum multipoint LOD score of 4.10 was obtained for 4 markers: D5S426, D5S493, D5S455, and D5S1998. Heterozygous carriers have significant prolongation of P-wave duration on ECGs compared with noncarriers (107 versus 85 ms on average; P=0.000012), but no differences between these 2 groups were detected for the PR interval, QRS complex, ST-segment duration, T-wave duration, QTc, and R-R interval (P>0.05). CONCLUSIONS: Our findings demonstrate that AF can be inherited as an autosomal recessive trait and define a novel genetic locus for AF on chromosome 5p13 (arAF1). A genetic link between AF and prolonged P-wave duration was identified. This study provides a framework for the ultimate cloning of the arAF1 gene, which will increase the understanding of the fundamental molecular mechanisms of atrial fibrillation.

Adolescent↗

Gene mining: a novel and powerful ensemble decision approach to hunting for disease genes using microarray expression profiling.

Current applications of microarrays focus on precise classification or discovery of biological types, for example tumor versus normal phenotypes in cancer research. Several challenging scientific tasks in the post-genomic epoch, like hunting for the genes underlying complex diseases from genome-wide gene expression profiles and thereby building the corresponding gene networks, are largely overlooked because of the lack of an efficient analysis approach. We have thus developed an innovative ensemble decision approach, which can efficiently perform multiple gene mining tasks. An application of this approach to analyze two publicly available data sets (colon data and leukemia data) identified 20 highly significant colon cancer genes and 23 highly significant molecular signatures for refining the acute leukemia phenotype, most of which have been verified either by biological experiments or by alternative analysis approaches. Furthermore, the globally optimal gene subsets identified by the novel approach have so far achieved the highest accuracy for classification of colon cancer tissue types. Establishment of this analysis strategy has offered the promise of advancing microarray technology as a means of deciphering the involved genetic complexities of complex diseases.

Acute Disease↗

Genomewide linkage scan identifies a novel susceptibility locus for restless legs syndrome on chromosome 9p.

Restless legs syndrome (RLS) is a common neurological disorder that affects 5%-12% of all whites. To genetically dissect this complex disease, we characterized 15 large and extended multiplex pedigrees, consisting of 453 subjects (134 affected with RLS). A familial aggregation analysis was performed, and SAGE FCOR was used to quantify the total genetic contribution in these families. A weighted average correlation of 0.17 between first-degree relatives was obtained, and heritability was estimated to be 0.60 for all types of relative pairs, indicating that RLS is a highly heritable trait in this ascertained cohort. A genomewide linkage scan, which involved >400 10-cM-spaced markers and spanned the entire human genome, was then performed for 144 individuals in the cohort. Model-free linkage analysis identified one novel significant RLS-susceptibility locus on chromosome 9p24-22 with a multipoint nonparametric linkage (NPL) score of 3.22. Suggestive evidence of linkage was found on chromosome 3q26.31 (NPL score 2.03), chromosome 4q31.21 (NPL score 2.28), chromosome 5p13.3 (NPL score 2.68), and chromosome 6p22.3 (NPL score 2.06). Model-based linkage analysis, with the assumption of an autosomal-dominant mode of inheritance, validated the 9p24-22 linkage to RLS in two families (two-point LOD score of 3.77; multipoint LOD score of 3.91). Further fine mapping confirmed the linkage result and defined this novel RLS disease locus to a critical interval. This study establishes RLS as a highly heritable trait, identifies a novel genetic locus for RLS, and will facilitate further cloning and identification of the genes for RLS.

Chromosome Mapping↗

Identification of an angiogenic factor that when mutated causes susceptibility to Klippel-Trenaunay syndrome.

Angiogenic factors are critical to the initiation of angiogenesis and maintenance of the vascular network. Here we use human genetics as an approach to identify an angiogenic factor, VG5Q, and further define two genetic defects of VG5Q in patients with the vascular disease Klippel-Trenaunay syndrome (KTS). One mutation is chromosomal translocation t(5;11), which increases VG5Q transcription. The second is mutation E133K identified in five KTS patients, but not in 200 matched controls. VG5Q protein acts as a potent angiogenic factor in promoting angiogenesis, and suppression of VG5Q expression inhibits vessel formation. E133K is a functional mutation that substantially enhances the angiogenic effect of VG5Q. VG5Q shows strong expression in blood vessels and is secreted as vessel formation is initiated. VG5Q can bind to endothelial cells and promote cell proliferation, suggesting that it may act in an autocrine fashion. We also demonstrate a direct interaction of VG5Q with another secreted angiogenic factor, TWEAK (also known as TNFSF12). These results define VG5Q as an angiogenic factor, establish VG5Q as a susceptibility gene for KTS, and show that increased angiogenesis is a molecular pathogenic mechanism of KTS.

Amino Acid Sequence↗

Premature myocardial infarction novel susceptibility locus on chromosome 1P34-36 identified by genomewide linkage analysis.

The most frequent causes of death and disability in the Western world are atherosclerotic coronary artery disease (CAD) and acute myocardial infarction (MI). This common disease is thought to have a polygenic basis with a complex interaction with environmental factors. Here, we report results of a genomewide search for susceptibility genes for MI in a well-characterized U.S. cohort consisting of 1,613 individuals in 428 multiplex families with familial premature CAD and MI: 712 with MI, 974 with CAD, and average age of onset of 44.4+/-9.7 years. Genotyping was performed at the National Heart, Lung, and Blood Institute Mammalian Genotyping Facility through use of 408 markers that span the entire human genome every 10 cM. Linkage analysis was performed with the modified Haseman-Elston regression model through use of the SIBPAL program. Three genomewide scans were conducted: single-point, multipoint, and multipoint performed on of white pedigrees only (92% of the cohort). One novel significant susceptibility locus was detected for MI on chromosomal region 1p34-36, with a multipoint allele-sharing P value of <10(-12) (LOD=11.68). Validation by use of a permutation test yielded a pointwise empirical P value of.00011 at this locus, which corresponds to a genomewide significance of P<.05. For the less restrictive phenotype of CAD, no genetic locus was detected, suggesting that CAD and MI may not share all susceptibility genes. The present study thus identifies a novel genetic-susceptibility locus for MI and provides a framework for the ultimate cloning of a gene for the complex disease MI.

Adult↗

An ensemble method for gene discovery based on DNA microarray data.

The advent of DNA microarray technology has offered the promise of casting new insights onto deciphering secrets of life by monitoring activities of thousands of genes simultaneously. Current analyses of microarray data focus on precise classification of biological types, for example, tumor versus normal tissues. A further scientific challenging task is to extract disease-relevant genes from the bewildering amounts of raw data, which is one of the most critical themes in the post-genomic era, but it is generally ignored due to lack of an efficient approach. In this paper, we present a novel ensemble method for gene extraction that can be tailored to fulfill multiple biological tasks including (i) precise classification of biological types; (ii) disease gene mining; and (iii) target-driven gene networking. We also give a numerical application for (i) and (ii) using a public microarrary data set and set aside a separate paper to address (iii).

Algorithms↗

Genetic linkage analysis of longitudinal hypertension phenotypes using three summary measures.

BACKGROUND: Longitudinal data often have multiple (repeated) measures recorded along a time trajectory. For example, the two cohorts from the Framingham Heart Study (GAW13 Problem 1) contain 21 and 5 repeated measures for hypertension phenotypes as well as epidemiological risk factors, respectively. Direct modelling of a large number of serially and biologically correlated traits in the context of linkage analysis can be prohibitively complex. Alternatively, we may consider using univariate transformation for linkage analysis of longitudinal repeated measures. RESULTS: We evaluated the utility of three conventional summary measures (mean, slope, and principal components) for genetic linkage analysis of longitudinal phenotypes by analyzing the chromosome 10 data of the Framingham Heart Study. Except for the temporal slope, all of the summary methods and the multivariate analysis identified the previously reported region, marker GATA64A09, for systolic blood pressure or high blood pressure. Further analysis revealed that this region may harbor gene(s) affecting human blood pressure at multiple stages of life. CONCLUSION: We conclude that mean and principal components are feasible alternatives for genetic linkage analysis of longitudinal phenotypes, but the slope might have a separate genetic basis from that of the original longitudinal phenotypes.

Adult Children↗

Multivariate sib-pair linkage analysis of longitudinal phenotypes by three step-wise analysis approaches.

BACKGROUND: Current statistical methods for sib-pair linkage analysis of complex diseases include linear models, generalized linear models, and novel data mining techniques. The purpose of this study was to further investigate the utility and properties of a novel pattern recognition technique (step-wise discriminant analysis) using the chromosome 10 linkage data from the Framingham Heart Study and by comparing it with step-wise logistic regression and linear regression. RESULTS: The three step-wise approaches were compared in terms of statistical significance and gene localization. Step-wise discriminant linkage analysis approach performed best; next was step-wise logistic regression; and step-wise linear regression was the least efficient because it ignored the categorical nature of disease phenotypes. Nevertheless, all three methods successfully identified the previously reported chromosomal region linked to human hypertension, marker GATA64A09. We also explored the possibility of using the discriminant analysis to detect gene x gene and gene x environment interactions. There was evidence to suggest the existence of gene x environment interactions between markers GATA64A09 or GATA115E01 and hypertension treatment and gene x gene interactions between markers GATA64A09 and GATA115E01. Finally, we answered the theoretical question "Is a trichotomous phenotype more efficient than a binary?" Unlike logistic regression, discriminant sib-pair linkage analysis might have more power to detect linkage to a binary phenotype than a trichotomous one. CONCLUSION: We confirmed our previous speculation that step-wise discriminant analysis is useful for genetic mapping of complex diseases. This analysis also supported the possibility of the pattern recognition technique for investigating gene x gene or gene x environment interactions.

Adult Children↗

Proteomic approach to coronary atherosclerosis shows ferritin light chain as a significant marker: evidence consistent with iron hypothesis in atherosclerosis.

Coronary artery disease (CAD) is the leading cause of mortality and morbidity in developed nations. We hypothesized that CAD is associated with distinct patterns of protein expression in the coronary arteries, and we have begun to employ proteomics to identify differentially expressed proteins in diseased coronary arteries. Two-dimensional (2-D) gel electrophoresis of proteins and subsequent mass spectrometric analysis identified the ferritin light chain as differentially expressed between 10 coronary arteries from patients with CAD and 7 coronary arteries from normal individuals. Western blot analysis indicated significantly increased expression of the ferritin light chain in the diseased coronary arteries (1.41 vs. 0.75; P = 0.01). Quantitative real-time PCR analysis showed that expression of ferritin light chain mRNA was decreased in diseased tissues (0.70 vs. 1.17; P = 0.013), suggesting that increased expression of ferritin light chain in CAD coronary arteries may be related to increased protein stability or upregulation of expression at the posttranscriptional level in the diseased tissues. Ferritin light chain protein mediates storage of iron in cells. We speculate that increased expression of the ferritin light chain may contribute to pathogenesis of CAD by modulating oxidation of lipids within the vessel wall through the generation of reactive oxygen species. Our results provide in situ proteomic evidence consistent with the "iron hypothesis," which proposes an association between excessive iron storage and a high risk of CAD. However, it is also possible that the increased ferritin expression in diseased coronary arteries is a consequence, rather than a cause, of CAD.

Adult↗

Longitudinal data analysis in pedigree studies.

Longitudinal family studies provide a valuable resource for investigating genetic and environmental factors that influence long-term averages and changes over time in a complex trait. This paper summarizes 13 contributions to Genetic Analysis Workshop 13, which include a wide range of methods for genetic analysis of longitudinal data in families. The methods can be grouped into two basic approaches: 1) two-step modeling, in which repeated observations are first reduced to one summary statistic per subject (e.g., a mean or slope), after which this statistic is used in a standard genetic analysis, or 2) joint modeling, in which genetic and longitudinal model parameters are estimated simultaneously in a single analysis. In applications to Framingham Heart Study data, contributors collectively reported evidence for genes that affected trait mean on chromosomes 1, 2, 3, 5, 8, 9, 10, 13, and 17, but most did not find genes affecting slope. Applications to simulated data suggested that even for a gene that only affected slope, use of a mean-type statistic could provide greater power than a slope-type statistic for detecting that gene. We report on the results of a small experiment that sheds some light on this apparently paradoxical finding, and indicate how one might form a more powerful test for finding a slope-affecting gene. Several areas for future research are discussed.

Cardiovascular Diseases↗