International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1722278

1722278 Vol 10 · Issue 2 Download Paper

Displaced and Suppressed: A Root-Cause Taxonomy of Test Failures in a Large Industrial Playwright/BDD Suite

Soniya V

Subject area: Science,Engineering and Technology  ·  Area of research: Software Testing

DOI: https://doi.org/10.64388/IREV10I2-1722278

Abstract

End-to-end (E2E) user-interface test suites at industrial scale are widely reported to make flakiness their dominant maintenance cost. Less widely examined is a second problem that compounds it: the site at which a failure is reported often diverges from the site of its true defect, and a further, under-examined subset of defects suppresses the failure signal entirely rather than merely displacing it. This paper reports an industrial case study of a 48-feature, approximately 852-scenario Playwright and Cucumber behaviour-driven test suite exercising a commercial multi-tenant software-as-a-service platform across three environments. Through practitioner-embedded root-cause analysis, latency instrumentation of the application under test, and codebase-wide pattern search, a taxonomy of seven recurring defect classes is derived and organized along a displacement/suppression axis. Five classes cause a correct expectation about test behaviour to fail at a location distant from the true cause; two classes cause a broken test to report a pass while verifying nothing. Measured latencies that motivate specific fixes are reported, the propagation of one defect class across eighteen page-object classes is quantified, and a dependency-partitioned parallel execution architecture adopted as a mitigating intervention for suite run time is described. The taxonomy is positioned against the flaky-test and test-smell literature, noting where the suppression classes are structurally analogous to previously reported “rotten green test” phenomena in unit testing, and the threats to validity that follow from a single-suite, single-team study are set out.

Keywords

Behaviour-Driven Development, End-To-End Testing, Flaky Tests, Silent Test Failure, Test Smells.

References

[1] Q. Luo, F. Hariri, L. Eloussi, and D. Marinov, “An empirical analysis of flaky tests,” in Proc. 22nd ACM SIGSOFT Int. Symp. Foundations of Software Engineering (FSE), 2014, pp. 643–653.

[2] M. Eck, F. Palomba, M. Castelluccio, and A. Bacchelli, “Understanding flaky tests: The developer's perspective,” in Proc. 27th ACM Joint European Software Engineering Conf. and Symp. on the Foundations of Software Engineering (ESEC/FSE), 2019, pp. 830–840.

[3] S. Zhang, D. Jalali, J. Wuttke, K. Muşlu, W. Lam, M. D. Ernst, and D. Notkin, “Empirically revisiting the test independence assumption,” in Proc. Int. Symp. on Software Testing and Analysis (ISSTA), 2014, pp. 385–396.

[4] M. Gruber, S. Lukasczyk, F. Kroiß, and G. Fraser, “An empirical study of flaky tests in Python,” in Proc. IEEE Conf. on Software Testing, Verification and Validation (ICST), 2021, pp. 148–158.

[5] W. Lam, P. Godefroid, S. Nath, A. Santhiar, and S. Thummalapenta, “Root causing flaky tests in a large-scale industrial setting,” in Proc. Int. Symp. on Software Testing and Analysis (ISSTA), 2019, pp. 204–215.

[6] A. van Deursen, L. Moonen, A. van den Bergh, and G. Kok, “Refactoring test code,” in Proc. 2nd Int. Conf. on Extreme Programming and Flexible Processes in Software Engineering (XP2001), 2001, pp. 92–95.

[7] J. Delplanque, A. Etien, S. Ducasse, and C. Fuhrman, “Rotten green tests,” in Proc. 41st Int. Conf. on Software Engineering (ICSE), 2019, pp. 500–511.

[8] M. Martinez, A. Etien, S. Ducasse, and C. Fuhrman, “RTj: A Java framework for detecting and refactoring rotten green test cases,” arXiv:1912.07322, 2019.

[9] A. Berndt, T. Bach, and S. Baltes, “Flaky tests in a large industrial database management system: An empirical study of fixed issue reports for SAP HANA,” arXiv:2602.03556, 2026.

[10] R. Santana, L. Martins, T. Virgínio, L. Soares, H. Costa, and I. Machado, “Refactoring Assertion Roulette and Duplicate Assert test smells: a controlled experiment,” arXiv:2207.05539, 2022.

[11] C. Landin, S. Tahvili, H. Haggren, M. Längkvist, A. Muhammad, and A. Loutfi, “Cluster-based parallel testing using semantic analysis,” in Proc. 2nd IEEE Int. Conf. on Artificial Intelligence Testing (AITest), 2020, pp. 99–106.

How to cite this paper

Soniya V "Displaced and Suppressed: A Root-Cause Taxonomy of Test Failures in a Large Industrial Playwright/BDD Suite" Iconic Research And Engineering Journals Volume 10 Issue 2 2026 Page 1212-1218 https://doi.org/10.64388/IREV10I2-1722278
Soniya V "Displaced and Suppressed: A Root-Cause Taxonomy of Test Failures in a Large Industrial Playwright/BDD Suite" Iconic Research And Engineering Journals, vol. 10, no. 2, Aug. 2026, doi: https://doi.org/10.64388/IREV10I2-1722278
Soniya V (2026). Displaced and Suppressed: A Root-Cause Taxonomy of Test Failures in a Large Industrial Playwright/BDD Suite. Iconic Research And Engineering Journals, 10(2). doi: https://doi.org/10.64388/IREV10I2-1722278
Soniya V "Displaced and Suppressed: A Root-Cause Taxonomy of Test Failures in a Large Industrial Playwright/BDD Suite" Iconic Research And Engineering Journals, vol. 10, no. 2, Aug. 2026. Crossref, https://doi.org/10.64388/IREV10I2-1722278
@article{1722278,
      author = {Soniya V},
      title = {Displaced and Suppressed: A Root-Cause Taxonomy of Test Failures in a Large Industrial Playwright/BDD Suite},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {10},
      number = {2},
      pages = {1212-1218},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1722278.pdf},
      abstract = {End-to-end (E2E) user-interface test suites at industrial scale are widely reported to make flakiness their dominant maintenance cost. Less widely examined is a second problem that compounds it: the site at which a failure is reported often diverges from the site of its true defect, and a further, under-examined subset of defects suppresses the failure signal entirely rather than merely displacing it. This paper reports an industrial case study of a 48-feature, approximately 852-scenario Playwright and Cucumber behaviour-driven test suite exercising a commercial multi-tenant software-as-a-service platform across three environments. Through practitioner-embedded root-cause analysis, latency instrumentation of the application under test, and codebase-wide pattern search, a taxonomy of seven recurring defect classes is derived and organized along a displacement/suppression axis. Five classes cause a correct expectation about test behaviour to fail at a location distant from the true cause; two classes cause a broken test to report a pass while verifying nothing. Measured latencies that motivate specific fixes are reported, the propagation of one defect class across eighteen page-object classes is quantified, and a dependency-partitioned parallel execution architecture adopted as a mitigating intervention for suite run time is described. The taxonomy is positioned against the flaky-test and test-smell literature, noting where the suppression classes are structurally analogous to previously reported “rotten green test” phenomena in unit testing, and the threats to validity that follow from a single-suite, single-team study are set out.},
      keywords = {Behaviour-Driven Development, End-To-End Testing, Flaky Tests, Silent Test Failure, Test Smells.},
      month = {August},
      doi = {https://doi.org/10.64388/IREV10I2-1722278}
  }