International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1723665

1723665 Vol 10 · Issue 4 Download Paper

Calibration Before Complexity: Forecasting Compound-Hazard Exposure for U.S. Freight Systems from Public Multi-Source Data

Sunday Michael Oyebiyi

Subject area: Science,Engineering and Technology  ·  Area of research: Forecasting Compound-Hazard

Abstract

Freight systems face hazards that seldom arrive alone, yet predictive studies typically model weather, disaster, trade and financial shocks in isolation. We assemble a leakage-controlled state–day panel for the United States that aligns disaster signals (NOAA Storm Events and FEMA declarations), freight exposure (Commodity Flow Survey and Freight Analysis Framework version 5) and macro-geopolitical context (imports, the CBOE Volatility Index and the Geopolitical Risk index), and forecast whether a composite hazard shock index, a proxy for conditions under which freight disruption becomes likely, will exceed its training-period 90th percentile on the following day. Four learners were trained to December 2022 and evaluated once from January 2024 (22,617 and 11,296 state-days), with 2023 withheld. Discrimination was moderate and nearly indistinguishable across three models (ROC-AUC 0.725–0.733 for gradient boosting, logistic regression and histogram-based gradient boosting; 0.706 for random forest), but probabilistic accuracy diverged sharply. Only gradient boosting outperformed a constant base-rate forecast (Brier skill score ≈ 0.09); logistic regression and histogram-based boosting over-predicted risk at every level (skill ≈ −0.8 to −0.9). Event prevalence rose from 10% to 17.2% between windows. Ablations showed that hazard persistence carries most of the signal, freight exposure adds little, and a wider macroeconomic block reduces out-of-sample discrimination. At a seven-day horizon, logistic regression was the most stable learner. A newsvendor illustration shows how uncalibrated probabilities near a decision threshold translate into systematic over-stocking. The study provides a transparent public-data benchmark and shows that, for logistics risk screening, calibration and temporal validation matter more than model complexity.

Keywords

Compound hazards; Freight resilience; Probability calibration; Early warning; Machine learning; Distribution shift; Decision-focused evaluation

References

[1] Baker, S. R., Bloom, N., & Davis, S. J. (2016). Measuring economic policy uncertainty. The Quarterly Journal of Economics, 131(4), 1593–1636. Oxford Academic

[2] Baryannis, G., Validi, S., Dani, S., & Antoniou, G. (2019). Supply chain risk management and artificial intelligence: State of the art and future research directions. International Journal of Production Research, 57(7), 2179–2202. Taylor & Francis

[3] Benigno, G., di Giovanni, J., Groen, J. J. J., & Noble, A. I. (2022). The GSCPI: A new barometer of global supply chain pressures (Staff Report No. 1017). Federal Reserve Bank of New York. Federal Reserve Bank of New York

[4] Bergmeir, C., & Benítez, J. M. (2012). On the use of cross-validation for time series predictor evaluation. Information Sciences, 191, 192–213. ScienceDirect

[5] Bertsimas, D., & Kallus, N. (2020). From predictive to prescriptive analytics. Management Science, 66(3), 1025–1044. INFORMS

[6] Bloom, N. (2009). The impact of uncertainty shocks. Econometrica, 77(3), 623–685. Wiley

[7] Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. Springer

[8] Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3. American Meteorological Society

[9] Brintrup, A., Pak, J., Ratiney, D., Pearce, T., Wichmann, P., Woodall, P., & McFarlane, D. (2020). Supply chain data analytics for predicting supplier disruptions: A case study in complex asset manufacturing. International Journal of Production Research, 58(11), 3330–3341. Taylor & Francis

[10] Caldara, D., & Iacoviello, M. (2022). Measuring geopolitical risk. American Economic Review, 112(4), 1194–1225. American Economic Association

[11] Carvalho, V. M., Nirei, M., Saito, Y. U., & Tahbaz-Salehi, A. (2021). Supply chain disruptions: Evidence from the Great East Japan Earthquake. The Quarterly Journal of Economics, 136(2), 1255–1321. Oxford Academic

[12] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. ACM

[13] Chopra, S., & Sodhi, M. S. (2004). Managing risk to avoid supply-chain breakdown. MIT Sloan Management Review, 46(1), 53–61.

[14] Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., Beam, A. L., Van Calster, B., Ghassemi, M., Liu, X., Reitsma, J. B., van Smeden, M., Boulesteix, A.-L., Camaradou, J. C., Celi, L. A., Denaxas, S., Denniston, A. K., Glocker, B., Golub, R. M., Harvey, H., Heinze, G., … Logullo, P. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ, 385, e078378. BMJ

[15] Collins, G. S., Reitsma, J. B., Altman, D. G., & Moons, K. G. M. (2015). Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): The TRIPOD statement. BMJ, 350, g7594. BMJ

[16] Craighead, C. W., Blackhurst, J., Rungtusanatham, M. J., & Handfield, R. B. (2007). The severity of supply chain disruptions: Design characteristics and mitigation capabilities. Decision Sciences, 38(1), 131–156. Wiley

[17] Cutter, S. L., Boruff, B. J., & Shirley, W. L. (2003). Social vulnerability to environmental hazards. Social Science Quarterly, 84(2), 242–261. Wiley

[18] DeLong, E. R., DeLong, D. M., & Clarke-Pearson, D. L. (1988). Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach. Biometrics, 44(3), 837–845. Crossref

[19] Dietterich, T. G. (1998). Approximate statistical tests for comparing supervised classification learning algorithms. Neural Computation, 10(7), 1895–1923. MIT Press

[20] Efron, B., & Tibshirani, R. J. (1993). An introduction to the bootstrap. Chapman & Hall/CRC.

[21] Elmachtoub, A. N., & Grigas, P. (2022). Smart “predict, then optimize.” Management Science, 68(1), 9–26. INFORMS

[22] Ermagun, A., & Levinson, D. (2018). Spatiotemporal traffic forecasting: Review and proposed directions. Transport Reviews, 38(6), 786–814. Taylor & Francis

[23] Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861–874. ScienceDirect

[24] Federal Emergency Management Agency. (n.d.). OpenFEMA dataset: Disaster declarations summaries – v2 [Data set]. Retrieved April 20, 2026, from FEMA

[25] Federal Reserve Bank of St. Louis. (n.d.). Federal Reserve Economic Data (FRED) [Data set]. Retrieved April 20, 2026, from Federal Reserve Bank of St. Louis

[26] Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189–1232. Project Euclid

[27] Gall, M., Borden, K. A., & Cutter, S. L. (2009). When do losses count? Six fallacies of natural hazards loss data. Bulletin of the American Meteorological Society, 90(6), 799–809. American Meteorological Society

[28] Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378. Taylor & Francis

[29] Gössling, S., Neger, C., Steiger, R., & Bell, R. (2023). Weather, climate change, and transport: A review. Natural Hazards, 118(2), 1341–1360. Springer

[30] Grinsztajn, L., Oyallon, E., & Varoquaux, G. (2022). Why do tree-based models still outperform deep learning on typical tabular data? In Advances in Neural Information Processing Systems (Vol. 35, pp. 507–520).

[31] Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (PMLR Vol. 70, pp. 1321–1330).

[32] Hanley, J. A., & McNeil, B. J. (1982). The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology, 143(1), 29–36. Radiology

[33] He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21(9), 1263–1284. IEEE

[34] Hosseini, S., Ivanov, D., & Dolgui, A. (2019). Review of quantitative methods for supply chain resilience analysis. Transportation Research Part E: Logistics and Transportation Review, 125, 285–307. ScienceDirect

[35] Hyndman, R. J., & Athanasopoulos, G. (2021). Forecasting: Principles and practice (3rd ed.). OTexts. OTexts

[36] Ivanov, D. (2020). Predicting the impacts of epidemic outbreaks on global supply chains: A simulation-based analysis on the coronavirus outbreak (COVID-19/SARS-CoV-2) case. Transportation Research Part E: Logistics and Transportation Review, 136, 101922. ScienceDirect

[37] Ivanov, D., & Dolgui, A. (2020). Viability of intertwined supply networks: Extending the supply chain resilience angles towards survivability. A position paper motivated by COVID-19 outbreak. International Journal of Production Research, 58(10), 2904–2915. Taylor & Francis

[38] Kaufman, S., Rosset, S., Perlich, C., & Stitelman, O. (2012). Leakage in data mining: Formulation, detection, and avoidance. ACM Transactions on Knowledge Discovery from Data, 6(4), Article 15. ACM

[39] Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems (Vol. 30, pp. 3146–3154).

[40] Khouja, M. (1999). The single-period (news-vendor) problem: Literature review and suggestions for future research. Omega, 27(5), 537–553. ScienceDirect

[41] King, G., & Zeng, L. (2001). Logistic regression in rare events data. Political Analysis, 9(2), 137–163. Oxford Academic

[42] Lipton, Z. C., Wang, Y.-X., & Smola, A. J. (2018). Detecting and correcting for label shift with black box predictors. In Proceedings of the 35th International Conference on Machine Learning (PMLR Vol. 80, pp. 3122–3130).

[43] Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4765–4774).

[44] Moreno-Torres, J. G., Raeder, T., Alaiz-Rodríguez, R., Chawla, N. V., & Herrera, F. (2012). A unifying view on dataset shift in classification. Pattern Recognition, 45(1), 521–530. ScienceDirect

[45] National Centers for Environmental Information. (n.d.). Storm events database [Data set]. National Oceanic and Atmospheric Administration. Retrieved April 20, 2026, from NOAA

[46] Niculescu-Mizil, A., & Caruana, R. (2005). Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning (pp. 625–632). ACM. ACM

[47] Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.

[48] Pettit, T. J., Fiksel, J., & Croxton, K. L. (2010). Ensuring supply chain resilience: Development of a conceptual framework. Journal of Business Logistics, 31(1), 1–21. Wiley

[49] Robinson, W. S. (1950). Ecological correlations and the behavior of individuals. American Sociological Review, 15(3), 351–357. Crossref

[50] Rose, A. (2004). Defining and measuring economic resilience to disasters. Disaster Prevention and Management, 13(4), 307–314. Emerald

[51] Rose, A., & Wei, D. (2013). Estimating the economic consequences of a port shutdown: The special role of resilience. Economic Systems Research, 25(2), 212–232. Taylor & Francis

[52] Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206–215. Nature

[53] Saerens, M., Latinne, P., & Decaestecker, C. (2002). Adjusting the outputs of a classifier to new a priori probabilities: A simple procedure. Neural Computation, 14(1), 21–41. MIT Press

[54] Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3), e0118432. PLOS

[55] Sheffi, Y., & Rice, J. B., Jr. (2005). A supply chain view of the resilient enterprise. MIT Sloan Management Review, 47(1), 41–48.

[56] Snyder, L. V., Atan, Z., Peng, P., Rong, Y., Schmitt, A. J., & Sinsoysal, B. (2016). OR/MS models for supply chain disruptions: A review. IIE Transactions, 48(2), 89–109. Taylor & Francis

[57] Steyerberg, E. W., Vickers, A. J., Cook, N. R., Gerds, T., Gonen, M., Obuchowski, N., Pencina, M. J., & Kattan, M. W. (2010). Assessing the performance of prediction models: A framework for traditional and novel measures. Epidemiology, 21(1), 128–138. PubMed

[58] Tang, C. S. (2006). Perspectives in supply chain risk management. International Journal of Production Economics, 103(2), 451–488. ScienceDirect

[59] Tashman, L. J. (2000). Out-of-sample tests of forecasting accuracy: An analysis and review. International Journal of Forecasting, 16(4), 437–450. ScienceDirect

[60] Tukamuhabwa, B. R., Stevenson, M., Busby, J., & Zorzini, M. (2015). Supply chain resilience: Definition, review and theoretical foundations for further study. International Journal of Production Research, 53(18), 5592–5623. Taylor & Francis

[61] U.S. Bureau of Transportation Statistics. (n.d.). Freight Analysis Framework (FAF) [Data set]. Retrieved April 20, 2026, from U.S. Bureau of Transportation Statistics

[62] U.S. Census Bureau. (n.d.). Commodity Flow Survey (CFS) [Data set]. Retrieved April 20, 2026, from U.S. Census Bureau

[63] Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., & Steyerberg, E. W. (2019). Calibration: The Achilles heel of predictive analytics. BMC Medicine, 17, Article 230. BMC

[64] Verschuur, J., Koks, E. E., Li, S., & Hall, J. W. (2023). Multi-hazard risk to global port infrastructure and resulting trade and logistics losses. Communications Earth & Environment, 4, Article 5. Nature

[65] Vlahogianni, E. I., Karlaftis, M. G., & Golias, J. C. (2014). Short-term traffic forecasting: Where we are and where we’re going. Transportation Research Part C: Emerging Technologies, 43, 3–19. ScienceDirect

[66] Whaley, R. E. (2000). The investor fear gauge. The Journal of Portfolio Management, 26(3), 12–17. The Journal of Portfolio Management

[67] Wyrembek, M., Baryannis, G., & Brintrup, A. (2025). Causal machine learning for supply chain risk prediction and intervention planning. International Journal of Production Research, 63(15), 5629–5648. Taylor & Francis

[68] Zadrozny, B., & Elkan, C. (2002). Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 694–699). ACM. ACM

[69] Zheng, G., Kong, L., & Brintrup, A. (2023). Federated machine learning for privacy preserving, collective supply chain risk prediction. International Journal of Production Research, 61(23), 8115–8132. Taylor & Francis

[70] Zscheischler, J., Westra, S., van den Hurk, B. J. J. M., Seneviratne, S. I., Ward, P. J., Pitman, A., AghaKouchak, A., Bresch, D. N., Leonard, M., Wahl, T., & Zhang, X. (2018). Future climate risk from compound events. Nature Climate Change, 8(6), 469–477. Nature

How to cite this paper

Sunday Michael Oyebiyi "Calibration Before Complexity: Forecasting Compound-Hazard Exposure for U.S. Freight Systems from Public Multi-Source Data" Iconic Research And Engineering Journals Volume 10 Issue 4 2026 Page 671-690
Sunday Michael Oyebiyi "Calibration Before Complexity: Forecasting Compound-Hazard Exposure for U.S. Freight Systems from Public Multi-Source Data" Iconic Research And Engineering Journals, vol. 10, no. 4, Oct. 2026
Sunday Michael Oyebiyi (2026). Calibration Before Complexity: Forecasting Compound-Hazard Exposure for U.S. Freight Systems from Public Multi-Source Data. Iconic Research And Engineering Journals, 10(4).
Sunday Michael Oyebiyi "Calibration Before Complexity: Forecasting Compound-Hazard Exposure for U.S. Freight Systems from Public Multi-Source Data" Iconic Research And Engineering Journals, vol. 10, no. 4, Oct. 2026.
@article{1723665,
      author = {Sunday Michael Oyebiyi},
      title = {Calibration Before Complexity: Forecasting Compound-Hazard Exposure for U.S. Freight Systems from Public Multi-Source Data},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {10},
      number = {4},
      pages = {671-690},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1723665.pdf},
      abstract = {Freight systems face hazards that seldom arrive alone, yet predictive studies typically model weather, disaster, trade and financial shocks in isolation. We assemble a leakage-controlled state–day panel for the United States that aligns disaster signals (NOAA Storm Events and FEMA declarations), freight exposure (Commodity Flow Survey and Freight Analysis Framework version 5) and macro-geopolitical context (imports, the CBOE Volatility Index and the Geopolitical Risk index), and forecast whether a composite hazard shock index, a proxy for conditions under which freight disruption becomes likely, will exceed its training-period 90th percentile on the following day. Four learners were trained to December 2022 and evaluated once from January 2024 (22,617 and 11,296 state-days), with 2023 withheld. Discrimination was moderate and nearly indistinguishable across three models (ROC-AUC 0.725–0.733 for gradient boosting, logistic regression and histogram-based gradient boosting; 0.706 for random forest), but probabilistic accuracy diverged sharply. Only gradient boosting outperformed a constant base-rate forecast (Brier skill score ≈ 0.09); logistic regression and histogram-based boosting over-predicted risk at every level (skill ≈ −0.8 to −0.9). Event prevalence rose from 10% to 17.2% between windows. Ablations showed that hazard persistence carries most of the signal, freight exposure adds little, and a wider macroeconomic block reduces out-of-sample discrimination. At a seven-day horizon, logistic regression was the most stable learner. A newsvendor illustration shows how uncalibrated probabilities near a decision threshold translate into systematic over-stocking. The study provides a transparent public-data benchmark and shows that, for logistics risk screening, calibration and temporal validation matter more than model complexity.},
      keywords = {Compound hazards; Freight resilience; Probability calibration; Early warning; Machine learning; Distribution shift; Decision-focused evaluation},
      month = {October},
  }