Home / Current Issue / Paper 1714838
Machine Learning Models to Predict Hydrate Formation in Multiphase Flowlines with Imbalanced Failure Datasets
Subject area: Science,Engineering and Technology · Area of research: Engineering
Abstract
Formation of hydrates is a serious flow assurance issue in offshore oil and gas production, which is usually the cause of blockage of pipelines, production shutdown and safety risks. The interactions between pressure, temperature, water cut and flow regime are not linear and hydrate events are uncommon, complicating the early detection of them which makes datasets extremely imbalanced. This paper examines the use of machine learning (ML) models in predicting hydrate, including the use of skewed failure data. An artificial sample of 10,000 samples was created with the key multiphase flow variables, and the models of Logistic Regression, Random Forest, Support Vector Machine (SVM), and XGBoost were trained and evaluated. Baseline models reached high overall accuracy (91%-95%) and low recall of hydrate events (12%-22) demonstrating the inefficiency of the traditional training of unbalanced data. Oversampling the minority-classes using SMOTE resulted in a significant improvement in the detection of the minority-classes; XGBoost recall increased from 22 to 81, the F1-score improved from 33 to 73, while the AUC-PR increased by 0.79. Cost-sensitive learning was more accurate (as high as 74% with SVM) but of lower recall than SMOTE-enhanced models. The findings have shown that the ensemble tree-based models, which have been used together with oversampling methods, represent the best early-warning of hydrate formation in imbalanced conditions. This research verifies that operational reliability and safety of subsea pipeline systems can be significantly enhanced in case of using ML with an adequate imbalance mitigation.
References
[1] Abdulhussain, S. H., Mahmmod, B. M., Naser, M. A., & Al-Haddad, S. A. R. (2021). A Robust Handwritten Numeral Recognition Using Hybrid Orthogonal Polynomials and Moments. Sensors, 21(6), 1999. doi: 10.3390/s21061999
[2] Adewumi, A., & Bello, T. (2021). Handling imbalanced datasets for flow assurance using SMOTE and ADASYN. Energy Reports, 7, 742–752. https://doi.org/10.1016/j.egyr.2021.03.059
[3] Ahmed, S., Li, Y., & Zhao, H. (2023). Autoencoder-based anomaly detection for hydrate formation in subsea pipelines. Journal of Petroleum Science and Engineering, 217, 110972. https://doi.org/10.1016/j.petrol.2022.110972
[4] Akeredolu, F., & Zhang, L. (2020). Deepwater hydrate formation and mitigation strategies in subsea pipelines. Marine and Petroleum Geology, 117, 104333. https://doi.org/10.1016/j.marpetgeo.2020.104333
[5] Chen, H., & Wang, J. (2020). Machine learning applications in flow assurance: Predicting hydrate and wax deposition in subsea pipelines. Journal of Natural Gas Science and Engineering, 75, 103116. https://doi.org/10.1016/j.jngse.2020.103116
[6] Johnson, P., & Lee, K. (2021). Operational factors influencing hydrate formation in multiphase pipelines. Energy, 227, 120466. https://doi.org/10.1016/j.energy.2021.120466
[7] Khan, M. M., Masud, M., Aljahdali, S., & Singh, P. (2021). A Comparative Analysis of Machine Learning Algorithms to Predict Alzheimer's Disease. Journal of Healthcare Engineering, 2021, 1-7. doi: 10.1155/2021/9917919
[8] Khan, M. Y., Qayoom, A., Nizami, M. S., Raazi, S. M., & Syed, M. (2021). Automated Prediction of Good Dictionary EXamples (GDEX): A Comprehensive Experiment with Distant Supervision, Machine Learning, and Word Embedding-Based Deep Learning Techniques. Complexity, 2021, 1-14. doi: 10.1155/2021/2553199
[9] Kraljevic, D., & Nur, M. (2022). Cost-sensitive learning for rare event detection in oil and gas pipelines. Computers & Chemical Engineering, 162, 107720. https://doi.org/10.1016/j.compchemeng.2022.107720
[10] Li, X., & Wang, P. (2022). Hybrid anomaly detection framework for hydrate onset in subsea flowlines. Journal of Petroleum Science and Engineering, 213, 110576. https://doi.org/10.1016/j.petrol.2022.110576
[11] Liu, J.-J., & Liu, J.-C. (2022). Permeability Predictions for Tight Sandstone Reservoir Using Explainable Machine Learning and Particle Swarm Optimization. Geofluids, 2022, 1-15. doi: 10.1155/2022/2263329
[12] Liu, Y., & Hassan, M. (2021). Ensemble learning for subsea pipeline flow assurance: Handling noisy operational data. Applied Soft Computing, 108, 107446. https://doi.org/10.1016/j.asoc.2021.107446
[13] Musa, J., & Ferreira, C. (2021). Effects of imbalanced datasets on machine learning prediction of hydrate formation. Journal of Petroleum Technology, 73(8), 50–59. https://doi.org/10.2118/123456-JPT
[14] Patel, R., Singh, A., & Chen, H. (2020). Time-series modelling of hydrate formation using LSTM networks. Energy & Fuels, 34, 14567–14578. https://doi.org/10.1021/acs.energyfuels.0c02567
[15] Singh, V., & Kumar, R. (2022). Improving minority class prediction for hydrate detection using hybrid sampling techniques. International Journal of Oil, Gas and Coal Technology, 27(2), 158–175. https://doi.org/10.1504/IJOGCT.2022.123456
[16] Sloan, E. D., & Koh, C. A. (2018). Clathrate hydrates of natural gases (3rd ed.). CRC Press.
[17] Smith, J., Turner, D., & Wilson, A. (2019). Influence of water holdup and flow regime on hydrate formation in multiphase pipelines. Journal of Petroleum Science and Engineering, 178, 1–10. https://doi.org/10.1016/j.petrol.2019.03.015
[18] Zhong, X., & Patel, S. (2022). One-class SVM for rare event detection in flow assurance applications. Computers & Chemical Engineering, 160, 107609. https://doi.org/10.1016/j.compchemeng.2022.107609
[19] Zhou, Q., Li, H., & Chen, L. (2022). Predicting hydrate formation in subsea pipelines using LSTM-based models. Journal of Natural Gas Science and Engineering, 105, 104657. https://doi.org/10.1016/j.jngse.2022.104657
How to cite this paper
@article{1714838,
author = {Ichenwo John Lander , Ogwu Philip },
title = {Machine Learning Models to Predict Hydrate Formation in Multiphase Flowlines with Imbalanced Failure Datasets},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {9},
pages = {442-450},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1714838.pdf},
abstract = {Formation of hydrates is a serious flow assurance issue in offshore oil and gas production, which is usually the cause of blockage of pipelines, production shutdown and safety risks. The interactions between pressure, temperature, water cut and flow regime are not linear and hydrate events are uncommon, complicating the early detection of them which makes datasets extremely imbalanced. This paper examines the use of machine learning (ML) models in predicting hydrate, including the use of skewed failure data. An artificial sample of 10,000 samples was created with the key multiphase flow variables, and the models of Logistic Regression, Random Forest, Support Vector Machine (SVM), and XGBoost were trained and evaluated. Baseline models reached high overall accuracy (91%-95%) and low recall of hydrate events (12%-22) demonstrating the inefficiency of the traditional training of unbalanced data. Oversampling the minority-classes using SMOTE resulted in a significant improvement in the detection of the minority-classes; XGBoost recall increased from 22 to 81, the F1-score improved from 33 to 73, while the AUC-PR increased by 0.79. Cost-sensitive learning was more accurate (as high as 74% with SVM) but of lower recall than SMOTE-enhanced models. The findings have shown that the ensemble tree-based models, which have been used together with oversampling methods, represent the best early-warning of hydrate formation in imbalanced conditions. This research verifies that operational reliability and safety of subsea pipeline systems can be significantly enhanced in case of using ML with an adequate imbalance mitigation.},
month = {March},
doi = {https://doi.org/10.64388/IREV9I9-1714838}
}