Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

A single factor underlies the metabolic syndrome: a confirmatory factor analysis.

OBJECTIVE: Confirmatory factor analysis (CFA) was used to test the hypothesis that the components of the metabolic syndrome are manifestations of a single common factor. RESEARCH DESIGN AND METHODS: Three different datasets were used to test and validate the model. The Spanish and Mauritian studies included 207 men and 203 women and 1,411 men and 1,650 women, respectively. A third analytical dataset including 847 men was obtained from a previously published CFA of a U.S. population. The one-factor model included the metabolic syndrome core components (central obesity, insulin resistance, blood pressure, and lipid measurements). We also tested an expanded one-factor model that included uric acid and leptin levels. Finally, we used CFA to compare the goodness of fit of one-factor models with the fit of two previously published four-factor models. RESULTS: The simplest one-factor model showed the best goodness-of-fit indexes (comparative fit index 1, root mean-square error of approximation 0.00). Comparisons of one-factor with four-factor models in the three datasets favored the one-factor model structure. The selection of variables to represent the different metabolic syndrome components and model specification explained why previous exploratory and confirmatory factor analysis, respectively, failed to identify a single factor for the metabolic syndrome. CONCLUSIONS: These analyses support the current clinical definition of the metabolic syndrome, as well as the existence of a single factor that links all of the core components.

Blood Pressure↗

Use of infrared spectroscopy for diagnosis of traumatic arthritis in horses.

OBJECTIVE: To evaluate use of infrared spectroscopy for diagnosis of traumatic arthritis in horses. ANIMALS: 48 horses with traumatic arthritis and 5 clinically and radiographically normal horses. PROCEDURES: Synovial fluid samples were collected from 77 joints in 48 horses with traumatic arthritis. Paired samples (affected and control joints) from 29 horses and independent samples from an affected (n = 12) or control (7) joint from 19 horses were collected for model calibration. A second set of 20 normal validation samples was collected from 5 clinically and radiographically normal horses. Fourier transform infrared spectra of synovial fluids were acquired and manipulated, and data from affected joints were compared with controls to identify spectroscopic features that differed significantly between groups. A classification model that used linear discriminant analysis was developed. Performance of the model was determined by use of the 2 validation datasets. RESULTS: A classification model based on 3 infrared regions classified spectra from the calibration dataset with overall accuracy of 97% (sensitivity, 93%; specificity, 100%). The model, with cost-adjusted prior probabilities of 0.60:0.40, yielded overall accuracy of 89% (sensitivity, 83%; specificity, 100%) for the first validation sample dataset and 100% correct classification of the second set of independent normal control joints. CONCLUSIONS AND CLINICAL RELEVANCE: The infrared spectroscopic patterns of fluid from joints with traumatic arthritis differed significantly from the corresponding patterns for controls. These alterations in absorption patterns may be used via an appropriate classification algorithm to differentiate the spectra of affected joints from those of controls.

Animals↗

Comparative factor analysis models for an empirical study of EEG data.

This paper (the first in a series) applies new, empirical factor analysis methods to the problem of "banding" EEG power spectra. A measure is introduced for the comparison of factor analysis results (factor loading matrices). The measure, Ambient Matrix Coherence (AC) is "geometrically unbiased", and invariant of so-called "oblique rotations." AC is used in a "stability computation" to find the dimension of the stable factor analysis solution common to several subsets of a given dataset. If the factor analysis model is appropriate, then the correct number of factors is empirically determined in this way. Stability computations were first performed on various simulated datasets to establish the robustness and efficacy of this method (for various noise levels). These techniques were then applied to EEG power spectra datasets for each of 8 leads. Comparison of these results indicated 3 stable factors in common to all 8 leads, an additional less stable factor in common to 5 leads, and weak stability for the six-dimensional solution for one lead.

Adult↗

Development and evaluation of models to predict the feed intake of dairy cows in early lactation.

Inaccurate prediction of dry matter intake (DMI) limits the ability of current models to anticipate the technical and economic consequences of adopting different strategies for production management on individual dairy farms. The objective of the present study was to develop an accurate, robust, and broadly applicable prediction model and to compare it with the current NRC model for dairy cows in early lactation. Among various functions, an exponential model was selected for its best fit to DMI data of dairy cows in early lactation. Daily DMI data (n = 8,547) for 3 groups of Holstein cows (at Illinois, New Hampshire, and Pennsylvania) were used in this study. Cows at Illinois and New Hampshire were fed totally mixed diets for the first 70 d of lactation. At Pennsylvania, data were for the first 63 d postpartum. Data from Illinois cows were used as the developmental dataset, and the other 2 datasets were used for model evaluation and validation. Data for BW, milk yield, and milk composition were only available for Illinois and New Hampshire cows; therefore, only these 2 datasets were used for model comparisons. The exponential model, fitted to the individual cow daily DMI data, explained an average of 74% of the total variation in daily DMI for Illinois data, 49% of the variation for New Hampshire data, 67% of the variation for Pennsylvania data, and 64% of the variation overall. Based on all model selection criteria used in this study, the exponential model for prediction of weekly DMI of individual cows was superior to the current NRC equation. The exponential model explained 85% of the variation in weekly mean DMI compared with 42% for the NRC equation. Compared with the relative prediction error of 6% for the exponential model, that associated with prediction using the NRC equation was 14%. The overall mean square prediction error value for individual cows was 5-fold higher for the NRC equation than for the exponential model (10.4 vs. 2.0 kg2/d2). The consistently accurate and robust prediction of DMI by the exponential model for all data-sets suggested that it could safely be used for predicting DMI in many circumstances.

Animal Nutritional Physiological Phenomena↗

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar↗

A comprehensive meta-analysis of tissue resident memory T cells and their roles in shaping immune microenvironment and patient prognosis in non-small cell lung cancer.

Tissue-resident memory T cells (TRM) are a specialized subset of long-lived memory T cells that reside in peripheral tissues. However, the impact of TRM-related immunosurveillance on the tumor-immune microenvironment (TIME) and tumor progression across various non-small-cell lung cancer (NSCLC) patient populations is yet to be elucidated. Our comprehensive analysis of multiple independent single-cell and bulk RNA-seq datasets of patient NSCLC samples generated reliable, unique TRM signatures, through which we inferred the abundance of TRM in NSCLC. We discovered that TRM abundance is consistently positively correlated with CD4+ T helper 1 cells, M1 macrophages, and resting dendritic cells in the TIME. In addition, TRM signatures are strongly associated with immune checkpoint and stimulatory genes and the prognosis of NSCLC patients. A TRM-based machine learning model to predict patient survival was validated and an 18-gene risk score was further developed to effectively stratify patients into low-risk and high-risk categories, wherein patients with high-risk scores had significantly lower overall survival than patients with low-risk. The prognostic value of the risk score was independently validated by the Cancer Genome Atlas Program (TCGA) dataset and multiple independent NSCLC patient datasets. Notably, low-risk NSCLC patients with higher TRM infiltration exhibited enhanced T-cell immunity, nature killer cell activation, and other TIME immune responses related pathways, indicating a more active immune profile benefitting from immunotherapy. However, the TRM signature revealed low TRM abundance and a lack of prognostic association among lung squamous cell carcinoma patients in contrast to adenocarcinoma, indicating that the two NSCLC subtypes are driven by distinct TIMEs. Altogether, this study provides valuable insights into the complex interactions between TRM and TIME and their impact on NSCLC patient prognosis. The development of a simplified 18-gene risk score provides a practical prognostic marker for risk stratification.

Humans↗

Large Language Model and Knowledge Graph-Driven AJCC Staging of Prostate Cancer Using Pathology Reports.

Background/Objectives: To develop an automated American Joint Committee on Cancer (AJCC) staging system for radical prostatectomy pathology reports using large language model-based information extraction and knowledge graph validation. Methods: Pathology reports from 152 radical prostatectomy patients were used. Five additional parameters (Prostate-specific antigen (PSA) level, metastasis stage (M-stage), extraprostatic extension, seminal vesicle invasion, and perineural invasion) were extracted using GPT-4.1 with zero-shot prompting. A knowledge graph was constructed to model pathological relationships and implement rule-based AJCC staging with consistency validation. Information extraction performance was evaluated using a local open-source large language model (LLM) (Mistral-Small-3.2-24B-Instruct) across 16 parameters. The LLM-extracted information was integrated into the knowledge graph for automated AJCC staging classification and data consistency validation. The developed system was further validated using pathology reports from 88 radical prostatectomy patients in The Cancer Genome Atlas (TCGA) dataset. Results: Information extraction achieved an accuracy of 0.973 and an F1-score of 0.986 on the internal dataset, and 0.938 and 0.968, respectively, on external validation. AJCC staging classification showed macro-averaged F1-scores of 0.930 and 0.833 for the internal and external datasets, respectively. Knowledge graph-based validation detected data inconsistencies in 5 of 150 cases (3.3%). Conclusions: This study demonstrates the feasibility of automated AJCC staging through the integration of large language model information extraction and knowledge graph-based validation. The resulting system enables privacy-protected clinical decision support for cancer staging applications with extensibility to broader oncologic domains.

artificial intelligence↗

High-Density Genome-Wide Association Mapping Identifies Candidate Loci Associated with Maize Stalk Cell Wall Composition.

Maize (Zea mays L.) stalk cell wall composition is a key determinant of forage digestibility, lodging resistance, and biomass utilization efficiency. Although previous genome-wide association studies (GWAS) have identified loci associated with lignin (LIG), cellulose (CEL), and hemicellulose (HC), advances in genomic resources provide an opportunity to revisit existing phenotypic datasets at substantially higher resolution. Here, we re-analyzed a maize association panel consisting of 341 diverse inbred lines using an expanded genotype dataset containing 10.77 million SNPs, two derived compositional indices (CEL/HC and [LIG/(CEL + HC)], and six complementary GWAS models. Across all traits and models, we identified 855 unique significant SNPs associated with 579 candidate genes. Among the traits examined, LIG/(CEL + HC) yielded the greatest number of associations, suggesting that indices representing the relative balance among cell wall components may better capture the genetic architecture of cell wall composition than individual component measurements alone. Integration of multiple GWAS models with functional enrichment, haplotype, and selective sweep analyses prioritized three biologically relevant candidate genes encoding a MYB58 transcription factor, the glycosyltransferase Xt9, and a putative xyloglucan 6-xylosyltransferase. Haplotype analysis revealed significant effects of Xt9 and the xyloglucan 6-xylosyltransferase on cell wall composition, while selective sweep analysis identified Xt9 as a target of repeated selection during maize domestication, ecological adaptation, and modern breeding. Although these candidate genes provide promising targets for future investigation, the associations identified here are based on a single association panel and require functional and independent population validation. Collectively, our results demonstrate how high-density genotyping combined with complementary GWAS models can refine candidate associations and generate testable hypotheses from existing phenotypic datasets.

cell wall composition↗

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

Spatial transcriptomics (ST) profiles genome-wide gene expression while preserving the two-dimensional spatial context of mRNA molecules within tissue sections, enabling studies of tissue architecture and microenvironment-associated biology. However, ST analysis remains challenging because data import, quality control, integration, deconvolution, spatial statistics, and visualization often require multiple software environments and reproducible parameter choices. This protocol presents a practical computational workflow for public ST datasets in R, beginning with data acquisition and software setup and proceeding through Seurat-based data loading, quality control, normalization, multi-sample integration, clustering, and spatially variable gene analysis. The workflow then applies complementary deconvolution strategies, including reference-guided SPOTlight analysis and unsupervised STdeconvolve topic modeling, followed by Giotto-based spatial cell-cell communication analysis and interactive region-of-interest (ROI) selection using a custom Python Dash application. By emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints, the protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation.

Spatial Transcriptomics↗

Causal Relationship Between Ischemic Stroke and Vascular Dementia: A Mendelian Randomization Study.

Ischemic stroke (IS) is a major cause of disability and mortality worldwide, and vascular dementia (VaD) is a common dementia subtype associated with cerebrovascular injury. Observational studies have suggested a relationship between IS and VaD, but these studies are vulnerable to confounding and reverse causality. This protocol describes a reproducible two-sample Mendelian randomization (MR) workflow for evaluating the potential causal association between IS and VaD using publicly available genome-wide association study (GWAS) summary statistics. Genetic instruments associated with IS were extracted from a public GWAS dataset, and outcome associations for VaD were obtained from a public VaD GWAS dataset. The corresponding dataset IDs are provided in the Protocol section. After outcome matching and allele harmonization, 51 single-nucleotide polymorphisms (SNPs) were retained for the final MR analysis. The workflow includes instrumental variable selection, linkage disequilibrium clumping, allele harmonization, instrument strength assessment, inverse variance weighted (IVW) analysis, weighted median analysis, MR-Egger analysis, heterogeneity testing, horizontal pleiotropy assessment, and leave-one-out sensitivity analysis. In the representative analysis, the IVW method showed a positive association between genetically predicted IS and VaD risk, and the weighted median method yielded a directionally concordant result. The MR-Egger estimate was directionally consistent but did not reach statistical significance. Therefore, these findings should be interpreted as suggestive evidence of a possible causal effect, rather than definitive proof of causality. This protocol may help researchers apply a transparent and reproducible MR workflow to investigate cerebrovascular disease-related outcomes using public GWAS data.

Humans↗

A Computational Workflow for Prioritizing Microbial Metabolite-Associated Host Genes in Constipation-Predominant Irritable Bowel Syndrome.

No standardized computational pipeline exists for systematically prioritizing microbial metabolite-associated host genes and protein-ligand complexes from publicly available chemical, genomic, and structural databases. This article describes an eight-stage workflow that accepts a user-defined set of gut microbiota-derived metabolites and produces a ranked shortlist of candidate metabolite-associated host genes, enriched biological pathways, and structurally prioritized protein-ligand complexes for experimental follow-up. The pipeline integrates (i) chemoinformatic metabolite profiling; (ii) multi-database candidate target prediction using protein-chemical interaction and ligand-based target-prediction tool and a molecular docking program; (iii) differential gene expression analysis of publicly available transcriptomic data; (iv) target-differentially expressed gene overlap; (v) protein-protein interaction network construction and pathway enrichment; (vi) molecular docking with a molecular docking program; (vii) 200 ns molecular dynamics simulation using a molecular dynamics engine with a protein force field used for molecular dynamics simulations; and (viii) MM-PBSA binding free-energy estimation. As a worked example, nine gut microbiota-derived or microbiota-modified metabolites representing short-chain fatty acids, bile acids, tryptophan-derived metabolites, and urolithin A were processed using the public IBS-C rectal mucosal transcriptomic dataset GSE36701. The workflow ranked 17 unique predicted metabolite-associated genes that were differentially expressed in this dataset. Docking, molecular dynamics simulation, and MM-PBSA analyses structurally prioritized five metabolite-protein complexes: lithocholic acid-VDR, lithocholic acid-NR1H4/FXR, ursodeoxycholic acid-NR1H4/FXR, tryptamine-HTR2A (simulated in an explicit 1-Palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC) lipid bilayer), and urolithin A-CASP3. The protocol is designed to be adaptable to other metabolite sets, disease transcriptomic datasets, and target classes; all outputs are hypothesis-generating computational predictions that require independent transcriptomic replication, protein-level validation, and functional ligand-response assays before causal or therapeutic conclusions can be drawn.

Irritable Bowel Syndrome↗

New insights into diagnostic values and mechanisms of ferroptosis associated with immune infiltration in diabetic kidney disease.

The pathogenesis of diabetic kidney disease (DKD) is complex and closely related to ferroptosis and immune dysregulation, but the relevance is unclear. The present study investigates the potential mechanisms of ferroptosis-related genes (FRGs) in DKD and their relationship with the immune-inflammatory response. It searches for new diagnostic biomarkers to help diagnose and treat DKD. Four Gene Expression Omnibus (GEO) datasets, GSE30528, GSE30529 and GSE30122 as the test set, and GSE96804 for validation, were analyzed. FRGs were obtained from GeneCards, and 47 ferroptosis-related differentially expressed genes (FRDEGs) were identified by intersecting with DKD-related differentially expressed genes. Functional enrichment analyses, including Gene Ontology, Kyoto Encyclopedia of Genes and Genomes, Gene Set Enrichment Analysis and Gene Set Variation Analysis, revealed that these FRDEGs are primarily associated with ferroptosis, hypoxia response and immune inflammation. Subsequently, the weighted gene co-expression network analysis (WGCNA) was employed to expand the ferroptosis-related gene network, and intersection of the 47 FRDEGs with key WGCNA module genes yielded 10 key genes. Based on the 10 key genes, the least absolute shrinkage and selection operator and support vector machine algorithms identified three hub genes [chemokine ligand 5 (CCL5), forkhead box C1 (FOXC1) and lactotransferrin (LTF)] for DKD diagnosis. Receiver operating characteristic curves confirmed their diagnostic value, with FOXC1 and LTF validated in the independent dataset. Immune infiltration analysis via CIBERSORT revealed eight immune cell types with significantly different infiltration levels between the DKD and control group in the integrated GEO datasets. Notably, both LTF and CCL5 showed a significant positive correlation with gamma delta T cells (γδT). Quantitative PCR results confirmed differential expression of the three hub genes in the DKD group, with elevated expression observed in DKD mice following intervention with rosiglitazone and hyperoside.

bioinformatics analysis↗

Analysis of rhythmic variance--ANORVA. A new simple method for detecting rhythms in biological time series.

Cyclic variations of variables are ubiquitous in biomedical science. A number of methods for detecting rhythms have been developed, but they are often difficult to interpret. A simple procedure for detecting cyclic variations in biological time series and quantification of their probability is presented here. Analysis of rhythmic variance (ANORVA) is based on the premise that the variance in groups of data from rhythmic variables is low when a time distance of one period exists between the data entries. A detailed stepwise calculation is presented including data entry and preparation, variance calculating, and difference testing. An example for the application of the procedure is provided, and a real dataset of the number of papers published per day in January 2003 using selected keywords is compared to randomized datasets. Randomized datasets show no cyclic variations. The number of papers published daily, however, shows a clear and significant (p < 0.03) circaseptan (period of 7 days) rhythm, probably of social origin.

Analysis of Variance↗

The incidence and cost of adverse events in Victorian hospitals 2003-04.

OBJECTIVES: To determine the incidence of adverse events in patients admitted in the year 2003-04 to selected Victorian hospitals; to identify the main hospital-acquired diagnoses; and to estimate the cost of these complications to the Victorian and Australian health system. DESIGN: The patient-level costing dataset for major Victorian public hospitals, 1 July 2003-30 June 2004, was analysed for adverse events by identifying C-prefixed diagnosis codes denoting complications, preventable or otherwise, arising during the course of hospital treatment. The in-hospital cost of adverse events was estimated using linear regression modelling, adjusting for age and comorbidity. MAIN OUTCOME MEASURES: Cost of each patient admission ("admitted episode"), length of stay and mortality. RESULTS: During the designated timeframe, 979,834 admitted episodes were in the sample, of which 67,435 (6.88%) had at least one adverse event. Patients with adverse events stayed about 10 days longer and had over seven times the risk of in-hospital death than those without complications. After adjusting for age and comorbidity, the presence of an adverse event adds dollar 6826 to the cost of each admitted episode. The total cost of adverse events in this dataset in 2003-04 was dollar 460.311 million, representing 15.7% of the total expenditure on direct hospital costs, or an additional 18.6% of the total inpatient hospital budget. CONCLUSION: Adverse events are associated with significant costs. Administrative datasets are a cost-effective source of information that can be used for a range of clinical governance activities to prevent adverse events.

Adolescent↗

Nursing terminology: a comparison of the ICNP and the nursing intervention lexicon and taxonomy.

The purpose of this paper is to provide an overview of nursing terminology work done to date and to compare the labels and subsumed terms of the recent alpha version of the International Classification of Nursing Practice (ICNP) with the verb terms (n = 147) of the interventions in a dataset of interventions (n = 7292) categorized using the Nursing Intervention Lexicon and Taxonomy (NILT). Two estimates were used to evaluate the adequacy of the ICNP terms for representing intervention terminology. Term matches were done using the NILT categories most similar to the ICNP action types. The ICNP action type 'observing' (and its subsumed terms) was best, accounting for 20% of the NILT verbs. The ICNP label 'observing' (and its subsumed terms) ranked first in accounting for 69% of the interventions in the NILT categories (CND & CV). The remaining action types had many fewer matches for the verbs and interventions in the NILT dataset. Thus it is possible to conclude that a relatively small subset of the verb terms, and interventions in a dataset of natural language interventions categorized using NILT have been captured by the INCP alpha version of Axis A.

Humans↗

Risks of leukemia in Japanese atomic bomb survivors, in women treated for cervical cancer, and in patients treated for ankylosing spondylitis.

The dose-response relationship for radiation-induced leukemia was examined in a pooled analysis of three exposed populations: Japanese atomic bomb survivors, women treated for cervical cancer, and patients irradiated for ankylosing spondylitis. A total of 383 leukemias were observed among 283,139 study subjects. Considering all leukemias apart from chronic lymphocytic leukemia, the optimal relative risk model had a dose response with a purely quadratic term representing induction and an exponential term consistent with cell sterilization at high doses; the addition of a linear induction term did not improve the fit of the model. The relative risk decreased with increasing time since exposure and increasing attained age, and there were significant (P < 0.00001) differences in the parameters of the model between datasets. These differences were related in part to the significant differences (P = 0.003) between the models fitted to the three main radiogenic leukemia subtypes (acute myeloid leukemia, acute lymphocytic leukemia, chronic myeloid leukemia). When the three datasets were considered together but the analysis was repeated separately for the three leukemia subtypes, for each subtype the optimal model included quadratic and exponential terms in dose. For acute myeloid leukemia and chronic myeloid leukemia, there were reductions of relative risk with increasing time after exposure, whereas for acute lymphocytic leukemia the relative risk decreased with increasing attained age. For each leukemia subtype considered separately, there was no indication of a difference between the studies in the relative risk and its distribution as a function of dose, age and time (P > 0.10 for all three subtypes). The nonsignificant indications of differences between the three datasets when leukemia subtypes were considered separately may be explained by random variation, although a contribution from differences in exposure dose-rate regimens, inhomogeneous dose distribution within the bone marrow, inadequate adjustment forcell sterilization effects, or errors in dosimetry could have played a role.

Adolescent↗

Haptic rendering of isosurfaces directly from medical images.

Virtual environments for surgical training, planning and rehearsal have the potential to significantly enhance patient treatment and diagnosis. Haptic feedback devices provide forces to the physician through a manipulator, simulating palpation, scalpel cuts, or retraction of tissue. While haptics have been studied in other fields, medical applications of haptics remain in their infancy. We propose a method of haptically rendering isosurfaces (representing hard structures) directly from anatomical datasets rather than through traditional intermediate graphical representations such as polygons. Our algorithm determines an implicit surface representation of a volumetric isosurface on the fly, and renders the structure using standard haptic algorithms to calculate the forces felt by the user. This approach has the advantage of providing easy access to the rich volume dataset containing the actual anatomy. By relating Hounsfield units to density, haptic rendering has the ability to provide different resistances based on the tissue type being rendered. We developed and tested our algorithm using a quadric implicit surface, a common primitive in computer graphics, and have applied it to a variety of anatomical image datasets. Our paper describes the algorithm in sufficient detail to facilitate reproduction by others.

Algorithms↗

A new fast heuristic for computing the breakpoint phylogeny and experimental phylogenetic analyses of real and synthetic data.

The breakpoint phylogeny is an optimization problem proposed by Blanchette et al. for reconstructing evolutionary trees from gene order data. These same authors also developed and implemented BPAnalysis [3], a heuristic method (based upon solving many instances of the travelling salesman problem) for estimating the breakpoint phylogeny. We present a new heuristic for this purpose; although not polynomial-time, our heuristic is much faster in practice than BPAnalysis. We present and discuss the results of experimentation on synthetic datasets and on the flowering plant family Campanulaceae with three methods: our new method, BPAnalysis, and the neighbor-joining method [25] using several distance estimation techniques. Our preliminary results indicate that, on datasets with slow evolutionary rates and large numbers of genes in comparison with the number of taxa (genomes), all methods recover quite accurate reconstructions of the true evolutionary history (although BPAnalysis is too slow to be practical), but that on datasets where the rate of evolution is high relative to the number of genes, the accuracy of all three methods is poor.

Algorithms↗