Search PubMed⌕ Search

Biomedical subjects

Marianthi Markatou

Publications and source records attributed to Marianthi Markatou.

6 recordsLinked to original sources

A statistical methodology for analyzing co-occurrence data from a large sample.

Determining important associations among items in a large database is challenging due to multiple simultaneous hypotheses and the ability to select weak associations that are statistically but not clinically significant. The simple application of the chi2 test among all possible pairs of items results in mostly inappropriate associations surpassing the traditional (alpha=.05, chi2=3.94) threshold. One can choose a stricter threshold to find stronger associations, but the choice may be arbitrary. We combined the volume test of Diaconis and Efron with a p-value plot to select a more rigorous and less arbitrary threshold. The volume test adjusts the p-value of the chi2-statistic. A plot of adjusted p-values (1 - p versus N(p)), where N(p) is the number of test statistics with a p-value greater than p, should be linear if there are no true associations. The point where the plot deviates from a line can be used as a threshold. We used linear regression to select the threshold in a reproducible fashion. In one experiment, we found that the method selected a threshold similar to that previously obtained by manually reviewing associations.

Calibration↗

Machine learning and word sense disambiguation in the biomedical domain: design and evaluation issues.

BACKGROUND: Word sense disambiguation (WSD) is critical in the biomedical domain for improving the precision of natural language processing (NLP), text mining, and information retrieval systems because ambiguous words negatively impact accurate access to literature containing biomolecular entities, such as genes, proteins, cells, diseases, and other important entities. Automated techniques have been developed that address the WSD problem for a number of text processing situations, but the problem is still a challenging one. Supervised WSD machine learning (ML) methods have been applied in the biomedical domain and have shown promising results, but the results typically incorporate a number of confounding factors, and it is problematic to truly understand the effectiveness and generalizability of the methods because these factors interact with each other and affect the final results. Thus, there is a need to explicitly address the factors and to systematically quantify their effects on performance. RESULTS: Experiments were designed to measure the effect of "sample size" (i.e. size of the datasets), "sense distribution" (i.e. the distribution of the different meanings of the ambiguous word) and "degree of difficulty" (i.e. the measure of the distances between the meanings of the senses of an ambiguous word) on the performance of WSD classifiers. Support Vector Machine (SVM) classifiers were applied to an automatically generated data set containing four ambiguous biomedical abbreviations: BPD, BSA, PCA, and RSV, which were chosen because of varying degrees of differences in their respective senses. Results showed that: 1) increasing the sample size generally reduced the error rate, but this was limited mainly to well-separated senses (i.e. cases where the distances between the senses were large); in difficult cases an unusually large increase in sample size was needed to increase performance slightly, which was impractical, 2) the sense distribution did not have an effect on performance when the senses were separable, 3) when there was a majority sense of over 90%, the WSD classifier was not better than use of the simple majority sense, 4) error rates were proportional to the similarity of senses, and 5) there was no statistical difference between results when using a 5-fold or 10-fold cross-validation method. Other issues that impact performance are also enumerated. CONCLUSION: Several different independent aspects affect performance when using ML techniques for WSD. We found that combining them into one single result obscures understanding of the underlying methods. Although we studied only four abbreviations, we utilized a well-established statistical method that guarantees the results are likely to be generalizable for abbreviations with similar characteristics. The results of our experiments show that in order to understand the performance of these ML methods it is critical that papers report on the baseline performance, the distribution and sample size of the senses in the datasets, and the standard deviation or confidence intervals. In addition, papers should also characterize the difficulty of the WSD task, the WSD situations addressed and not addressed, as well as the ML methods and features used. This should lead to an improved understanding of the generalizablility and the limitations of the methodology.

Algorithms↗

Inter-patient distance metrics using SNOMED CT defining relationships.

BACKGROUND: Patient-based similarity metrics are important case-based reasoning tools which may assist with research and patient care applications. Ontology and information content principles may be potentially helpful tools for similarity metric development. METHODS: Patient cases from 1989 through 2003 from the Columbia University Medical Center data repository were converted to SNOMED CT concepts. Five metrics were implemented: (1) percent disagreement with data as an unstructured "bag of findings," (2) average links between concepts, (3) links weighted by information content with descendants, (4) links weighted by information content with term prevalence, and (5) path distance using descendants weighted by information content with descendants. Three physicians served as gold standard for 30 cases. RESULTS: Expert inter-rater reliability was 0.91, with rank correlations between 0.61 and 0.81, representing upper-bound performance. Expert performance compared to metrics resulted in correlations of 0.27, 0.29, 0.30, 0.30, and 0.30, respectively. Using SNOMED axis Clinical Findings alone increased correlation to 0.37. CONCLUSION: Ontology principles and information content provide useful information for similarity metrics but currently fall short of expert performance.

Algorithms↗

Pharmacodynamics of PEG-IFN alpha differentiate HIV/HCV coinfected sustained virological responders from nonresponders.

Pegylated interferon (PEG-IFN) has become standard therapy for hepatitis C virus (HCV) infection. We evaluated whether PEG-IFN pharmacodynamics and pharmacokinetics account for differences in treatment outcome and whether these parameters might be predictors of therapeutic outcome. Twenty-four IFN-naïve, HCV/human immunodeficiency virus-coinfected patients received PEG-IFN alpha-2b (1.5 microg/kg) once weekly plus daily ribavirin (1000 or 1200 mg) for up to 48 weeks. HCV RNA and PEG-IFN alpha concentrations were obtained from samples collected frequently after the first 3 PEG-IFN doses. We modeled HCV kinetics incorporating pharmacokinetic and pharmacodynamic parameters. Although PEG-IFN concentrations and pharmacokinetic parameters were similar in sustained virological responders (SVRs) and nonresponders (NRs), the PEG-IFN alpha-2b concentration that decreases HCV production by 50% (EC50) was lower in SVRs compared with NRs (0.04 vs. 0.45 microg/L [P = .014]). Additionally, the median therapeutic quotient (i.e., the ratio between average PEG-IFN concentration and EC50 [C/EC50]), and the PEG-IFN concentration at day 7 divided by EC50 (C(7)/EC50) were significantly increased in SVRs compared with NRs after the first (10.1 vs. 1.0 [P = .012], 2.8 vs. 0.3 [P = .007], respectively) and second (14.0 vs. 1.1 [P = .016], 5.4 vs. 0.4 [P = .02], respectively) PEG-IFN doses. All 3 parameters may be used to identify NRs. In conclusion, PEG-IFN concentrations and pharmacokinetic parameters do not differ between SVRs and NRs. In contrast, pharmacodynamic measurements-namely EC50, the therapeutic quotient, and C(7)/EC50--are different in coinfected SVRs and NRs. These parameters might be useful predictors of treatment outcome during the first month of therapy.

Antiviral Agents↗

Mining a clinical data warehouse to discover disease-finding associations using co-occurrence statistics.

This paper applies co-occurrence statistics to discover disease-finding associations in a clinical data warehouse. We used two methods, chi2 statistics and the proportion confidence interval (PCI) method, to measure the dependence of pairs of diseases and findings, and then used heuristic cutoff values for association selection. An intrinsic evaluation showed that 94 percent of disease-finding associations obtained by chi2 statistics and 76.8 percent obtained by the PCI method were true associations. The selected associations were used to construct knowledge bases of disease-finding relations (KB-chi2, KB-PCI). An extrinsic evaluation showed that both KB-chi2 and KB-PCI could assist in eliminating clinically non-informative and redundant findings from problem lists generated by our automated problem list summarization system.

Chi-Square Distribution↗

Virus dynamics and immune responses during treatment in patients coinfected with hepatitis C and HIV.

Mathematical modeling of the biological effect of interferon on virus decay permits the quantification of the efficacy (epsilon) of blocking virion production in different patient populations. The viral dynamic and immunologic responses of hepatitis C virus (HCV) infection to daily interferon therapy were characterized in twelve patients co-infected with human immunodeficiency virus (HIV). Three out of the twelve patients (25%) achieved an early viral response, a two-log reduction in HCV RNA by week 12. The mean epsilon of IFN-alpha in blocking HCV and HIV production were 72% and 74%, respectively. For HCV epsilon was highest (97%) in the one patient who had a sustained viral response, while it was reduced in the other two patients (68% and 77%). Baseline HCV RNA and the number of CD3+CD56+16+ cells were inversely related (r = -0.89, p = 0.03), and baseline HCV-specific immune responses were significantly higher in the three patients with 2-log viral load reductions. These data suggest that: 1) interferon efficacy at blocking virion production is correlated with treatment outcome in HIV/HCV co-infected patients, 2) that immunodeficient patients can respond to standard IFN-alpha, 3) that both innate and adaptive immune responses may be important determinants of HCV RNA decline in response to interferon.

Adult↗