Home / Current Issue / Paper 1720339
Artificial Intelligence-Based Prediction of Bioremediation of Crude Oil-Contaminated Soil
Subject area: Science,Engineering and Technology · Area of research: Machine Learning, Artificial Intelligence
Abstract
Crude oil contamination of soil remains a critical environmental challenge, particularly across petroleum-producing regions of the developing world. Although biological remediation strategies have attracted considerable research interest over the past decade, the empirical monitoring of remediation processes is constrained by the labor-intensive nature of laboratory measurements and the absence of reliable predictive frameworks. This work introduces a data-driven, machine-learning framework for forecasting bioremediation efficiency (BE%) in crude oil-contaminated soil, using a dataset of 14 monitored field samples measured fortnightly across an eight-week remediation cycle. To overcome the data scarcity challenge, a kinetic spline augmentation protocol expanded the original 56-observation dataset to 2,114 temporally dense records. Six supervised regression algorithms, namely Linear, Ridge, and Lasso Regression, Support Vector Regression, Gradient Boosting Regression, and a Multilayer Perceptron Artificial Neural Network (MLP-ANN), were trained and assessed using a five-fold cross-validation scheme. Among the six architectures tested, the MLP-ANN delivered the strongest predictive accuracy, returning a coefficient of determination (R²) of 0.9803, an adjusted R² of 0.9802, a root-mean-square error (RMSE) of 2.9613%, a mean absolute error (MAE) of 2.0970%, and a mean absolute percentage error (MAPE) of 4.29%, surpassing every other model evaluated. When deployed as a binary classifier at a regulatory compliance threshold of BE% of at least 70%, the ANN achieved a classification accuracy of 96.45%. SHapley Additive exPlanations (SHAP) analysis identified oil and grease concentration as the dominant predictor of remediation outcome, contributing a mean absolute SHAP value of 16.89, followed by moisture content (1.33), microbial count (1.11), and elapsed treatment time (0.39). Residual diagnostic analyses confirmed approximate normality and homoscedasticity of model errors. These findings establish the ANN as a robust, interpretable, and practically deployable tool for real-time monitoring and decision support in soil bioremediation programmes.
Keywords
bioremediation efficiency, artificial neural network, machine learning, crude oil contamination, SHAP interpretability, soil remediation prediction.
How to cite this paper
@article{1720339,
author = {Gaius Iliya, Abdulsalam Surajudeen, Kabiru Ibrahim Musa},
title = {Artificial Intelligence-Based Prediction of Bioremediation of Crude Oil-Contaminated Soil},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {2},
pages = {3073-3086},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1720339.pdf},
abstract = {Crude oil contamination of soil remains a critical environmental challenge, particularly across petroleum-producing regions of the developing world. Although biological remediation strategies have attracted considerable research interest over the past decade, the empirical monitoring of remediation processes is constrained by the labor-intensive nature of laboratory measurements and the absence of reliable predictive frameworks. This work introduces a data-driven, machine-learning framework for forecasting bioremediation efficiency (BE%) in crude oil-contaminated soil, using a dataset of 14 monitored field samples measured fortnightly across an eight-week remediation cycle. To overcome the data scarcity challenge, a kinetic spline augmentation protocol expanded the original 56-observation dataset to 2,114 temporally dense records. Six supervised regression algorithms, namely Linear, Ridge, and Lasso Regression, Support Vector Regression, Gradient Boosting Regression, and a Multilayer Perceptron Artificial Neural Network (MLP-ANN), were trained and assessed using a five-fold cross-validation scheme. Among the six architectures tested, the MLP-ANN delivered the strongest predictive accuracy, returning a coefficient of determination (R²) of 0.9803, an adjusted R² of 0.9802, a root-mean-square error (RMSE) of 2.9613%, a mean absolute error (MAE) of 2.0970%, and a mean absolute percentage error (MAPE) of 4.29%, surpassing every other model evaluated. When deployed as a binary classifier at a regulatory compliance threshold of BE% of at least 70%, the ANN achieved a classification accuracy of 96.45%. SHapley Additive exPlanations (SHAP) analysis identified oil and grease concentration as the dominant predictor of remediation outcome, contributing a mean absolute SHAP value of 16.89, followed by moisture content (1.33), microbial count (1.11), and elapsed treatment time (0.39). Residual diagnostic analyses confirmed approximate normality and homoscedasticity of model errors. These findings establish the ANN as a robust, interpretable, and practically deployable tool for real-time monitoring and decision support in soil bioremediation programmes.},
keywords = {bioremediation efficiency, artificial neural network, machine learning, crude oil contamination, SHAP interpretability, soil remediation prediction.},
month = {August},
}