Search PubMed⌕ Search

PubMed · 16408921

Improving the linearity of infrared diffuse reflection spectroscopy data for quantitative analysis: an application in quantifying organophosphorus contamination in soil.

Abstract

Diffuse reflection data are presented for ethyl methylphosphonate in a fine Utah dirt sample as a model system for organophosphate-contaminated soil. The data revealed a chemometric artifact when the spectra were represented in Kubelka-Munk units that manifests as a linear dependence of spectral peak height on variations in the observed baseline position (i.e., the position of the observed transmission intensity where no absorption features occur in the sample spectrum). We believe that this artifact is the result of the mathematical process by which the raw data are converted into Kubelka-Munk units, and we developed a numerical strategy for compensating for the observed effect and restoring chemometric precision to the diffuse reflection data for quantitative analysis while retaining the benefits of linear calibration afforded by the Kubelka-Munk approach. We validated our Kubelka-Munk correction strategy by repeating the experiment using a simpler system--pure caffeine in potassium bromide. The numerical preprocessing includes conventional multiplicative scatter correction coupled with a baseline offset correction that facilitates the use of quantitative diffuse reflection data in the Kubelka-Munk formalism for the quantitation of contaminants in a complex soil matrix, but is also applicable to more fundamental diffuse reflection quantitative analysis experiments.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alan C Samuels, Changjiang Zhu, Barry R Williams, Avishai Ben-David, Ronald W Miles, Melissa Hulet. 2006-01-15. Improving the linearity of infrared diffuse reflection spectroscopy data for quantitative analysis: an application in quantifying organophosphorus contamination in soil.. https://doi.org/10.1021/ac0509859

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data.

Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.

Linear Models↗

A novel QSPR model for predicting θ (lower critical solution temperature) in polymer solutions using molecular descriptors.

In this study, we present a new model that has been developed for the prediction of θ (lower critical solution temperature) using a database of 169 data points that include 12 polymers and 67 solvents. For the characterization of polymer and solvent molecules, a number of molecular descriptors (topological, physicochemical,steric and electronic) were examined. The best subset of descriptors was selected using the elimination selection-stepwise regression method. Multiple linear regression (MLR) served as the statistical tool to explore the potential correlation among the molecular descriptors and the experimental data. The prediction accuracy of the MLR model was tested using the leave-one-out cross validation procedure, validation through an external test set and the Y-randomization evaluation technique. The domain of applicability was finally determined to identify the reliable predictions.

Linear Models↗

Fit of four curve-linear models to decay profiles for pest control substances in soil.

Experiments that investigate the pattern of degradation of pest control substances in soil are often undertaken to estimate the persistence of compounds in the environment. Mathematical models are typically fit to decay data to facilitate the interpretation of the results and make predictions concerning the environmental fate of xenobiotics in soil. Four mathematical models were fit to 61 data sets to compare their performance in conforming to empirical patterns of degradation of pest control substances in soil. The use of composite residual plots allowed comparisons of the performance of the different models over many data sets. While an exponential model, estimated using nonlinear regression, fit many data sets very well, a shift-log, biexponential, and Monod equation appears superior in many cases, and systematic deviations from data sets are often less evident with the latter models. A knowledge of the patterns of bias typically exhibited by each model across many data sets may be useful for selecting models with reduced bias when fitting individual data sets.

Linear Models↗