International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1709872

1709872 Vol 9 · Issue 1 Download Paper

Design and Evaluation of an Anti-Plagiarism System Using Semantic Code Analysis

Idowu Olugbenga Adewumi Samuel Eleojo Agene Victoria Bola Oyekunle

Subject area: Science,Engineering and Technology  ·  Area of research: Software Engineering

Abstract

The predominance of source code plagiarism in educational and expert contexts has emphasized the boundaries of outdated recognition tools that depend heavily on syntactic similarity, such as string matching and token based assessments. This research work suggested and assesses a semantic code examination based anti-plagiarism system designed to ascertain three divergent types of plagiarism: Type I (superficial changes), Type II (structural modifications), and Type III (logic-preserving transformations). The system incorporates Abstract Syntax Tree (AST) demonstrations, graph based comparison metrics, and supervised machine learning representations to decode abysmal semantic connections amongst code samples. Assessment was piloted on a scraped dataset containing 100 Python code pairs, comprising both plagiarized and non-plagiarized samples. The projected system attained high ordering performance, with a macro averaged precision of 0.92, recall of 0.88, and F1-score of 0.90. AST-based investigation reliably outclassed etymological procedures, predominantly in identifying multifaceted plagiarism: for Type III cases, the semantic method yielded an F1-score of 0.86, matched to 0.55 for string matching methods. Between comparison metrics tested, Tree Edit Distance (TED) accomplished the maximum F1-score (0.93), whereas the joined metric vector presented a stable presentation across all classes (F1-score: 0.90). The Random Forest classifier established higher effectiveness above other machine learning and rule based prototypes, achieving a macro F1-score of 0.90, with a confusion matrix representing high true positive rates across all sessions. These outcomes asserted the effectiveness of semantic and structure aware techniques in discovering varied procedures of code plagiarism and highlighted the significance of incorporating graph theoretic methods with machine learning for robust taxonomy. The verdicts advocate for wider acceptance of semantic detection systems in educational technology, software forensics, and automated code review platforms.

Keywords

Code Plagiarism Detection, Abstract Syntax Tree (AST), Semantic Analysis, Graph Similarity, Machine Learning, Random Forest, Tree Edit Distance, Code Obfuscation, Multi-class Classification, Software Forensics

References

[1] Baxter, Glen, et al. 1998. “Clone Detection Using Abstract Syntax Trees.” Proceedings of the International Conference on Software Maintenance, 368–77.

[2] Jiang, Lingxiao, Ghassan Misherghi, Zhendong Su, and Stephane Glondu. 2007. “DECKARD: Scalable and Accurate TreeBased Detection of Code Clones.” In Proceedings of the 29th International Conference on Software Engineering (ICSE), 96–105. IEEE. arXiv+6ResearchGate+6InK at SMU+6

[3] Koschke, Rainer. 2007. “Survey of Research on Software Clones.” In Duplication, Redundancy, and Similarity in Software, edited by R. Koschke, E. Merlo, and A. Walenstein, Dagstuhl Seminar Proceedings 06301, 1–28. Schloss Dagstuhl, Germany. ResearchGate

[4] Prechelt, Lutz, Guido Malpohl, and Michael Philippsen. 2002. “Finding Plagiarisms among a Set of Programs with JPlag.” Journal of Universal Computer Science 8 (11): 1016–38. lib.jucs.org+2jucs.org+2jucs.org+2

[5] Schleimer, Saul, Daniel S. Wilkerson, and Alexander Aiken. 2003. “Winnowing: Local Algorithms for Document Fingerprinting.” In Proceedings of the 2003 ACM SIGMOD International Conference on Management of Data, San Diego, CA, June 9–12. ACM. ACM Digital Library+3ResearchGate+3ResearchGate+3

[6] Schleimer, Saul D., Daniel S. Wilkerson, and Alexander Aiken. 2003. Winnowing: Local Algorithms for Document Fingerprinting. UIC/CS Technical Report. ResearchGate

[7] Sajnani, Hitesh, Vaibhav Saini, Jeffrey Svajlenko, Chanchal Roy, and Cristina V. Lopes. 2016. “SourcererCC: Scalable Clone Detection at GitHub Scale.” Proceedings of the 2016 IEEE/ACM 38th International Conference on Software Engineering, 115–26. arXiv+1arXiv+1

[8] White, Mike, et al. 2016. “Deep Learning Code Fragments for Clone Detection.” ACM Transactions on Software Engineering and Methodology. (Use search to retrieve)

[9] Nguyen, Ha, and et al. 2012. “Detecting Semantic Code Clones Using Program Dependency Graphs.” Proceedings of the International Conference on Automated Software Engineering. (Use search to retrieve)

[10] Haldar, S., et al. 2012. “Detection of LogicBased Code Plagiarism.” Journal of Software Maintenance and Evolution: Research and Practice. (Use search)

[11] Guo, et al. 2017. “Learning to Detect Code Clones with Graph Neural Networks.” Proceedings of the ACM/IEEE International Conference on Software Engineering. (Use search)

[12] Lopes, Cristina V., et al. 2010. “DéjàVu: A Map of Code Duplicates on GitHub.” MSR (Mining Software Repositories).

[13] Alrabaee, S., et al. 2014. “ObfuscationResilient Code Plagiarism Detection.” Proceedings of the IEEE International Conference on Software Maintenance and Evolution.

[14] Krinke, Jürgen. 2001. “Identifying Similar Code with Program Dependence Graphs.” Proceedings of the International Conference on Software Engineering.

[15] Liu, C., et al. 2006. “GPLAG: Plagiarism Detection for Generic Programming Languages.” IEEE Transactions on Knowledge and Engineering.

[16] Wang, Wenhan, Ge Li, Bo Ma, Xin Xia, and Zhi Jin. 2020. “Detecting Code Clones with Graph Neural Network and FlowAugmented Abstract Syntax Tree.” arXiv. InK at SMU+3arXiv+3arXiv+3arXiv

[17] Ragkhitwetsagul, Chutipong, et al. 2018. “A Comprehensive Survey on Software Code Clones and Plagiarism.” Journal of Systems and Software.

[18] Roy, Chanchal K., and James R. Cordy. 2009. “A Survey on Software Clone Detection Research.” Queen’s University Technical Report.

[19] Svajlenko, Jeffrey, et al. 2014. “Evaluating Modern Clone Detection Tools.” Proceedings of the 2014 IEEE International Conference on Software Maintenance and Evolution.

How to cite this paper

Idowu Olugbenga Adewumi, Samuel Eleojo Agene, Victoria Bola Oyekunle "Design and Evaluation of an Anti-Plagiarism System Using Semantic Code Analysis" Iconic Research And Engineering Journals Volume 9 Issue 1 2025 Page 1891-1903
Idowu Olugbenga Adewumi, Samuel Eleojo Agene, Victoria Bola Oyekunle "Design and Evaluation of an Anti-Plagiarism System Using Semantic Code Analysis" Iconic Research And Engineering Journals, vol. 9, no. 1, Jul. 2025
Idowu Olugbenga Adewumi, Samuel Eleojo Agene, Victoria Bola Oyekunle (2025). Design and Evaluation of an Anti-Plagiarism System Using Semantic Code Analysis. Iconic Research And Engineering Journals, 9(1).
Idowu Olugbenga Adewumi, Samuel Eleojo Agene, Victoria Bola Oyekunle "Design and Evaluation of an Anti-Plagiarism System Using Semantic Code Analysis" Iconic Research And Engineering Journals, vol. 9, no. 1, Jul. 2025.
@article{1709872,
      author = {Idowu Olugbenga Adewumi, Samuel Eleojo Agene, Victoria Bola Oyekunle},
      title = {Design and Evaluation of an Anti-Plagiarism System Using Semantic Code Analysis},
      journal = {Iconic Research And Engineering Journals},
      year = {2025},
      volume = {9},
      number = {1},
      pages = {1891-1903},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1709872.pdf},
      abstract = {The predominance of source code plagiarism in educational and expert contexts has emphasized the boundaries of outdated recognition tools that depend heavily on syntactic similarity, such as string matching and token based assessments. This research work suggested and assesses a semantic code examination based anti-plagiarism system designed to ascertain three divergent types of plagiarism: Type I (superficial changes), Type II (structural modifications), and Type III (logic-preserving transformations). The system incorporates Abstract Syntax Tree (AST) demonstrations, graph based comparison metrics, and supervised machine learning representations to decode abysmal semantic connections amongst code samples. Assessment was piloted on a scraped dataset containing 100 Python code pairs, comprising both plagiarized and non-plagiarized samples. The projected system attained high ordering performance, with a macro averaged precision of 0.92, recall of 0.88, and F1-score of 0.90. AST-based investigation reliably outclassed etymological procedures, predominantly in identifying multifaceted plagiarism: for Type III cases, the semantic method yielded an F1-score of 0.86, matched to 0.55 for string matching methods. Between comparison metrics tested, Tree Edit Distance (TED) accomplished the maximum F1-score (0.93), whereas the joined metric vector presented a stable presentation across all classes (F1-score: 0.90). The Random Forest classifier established higher effectiveness above other machine learning and rule based prototypes, achieving a macro F1-score of 0.90, with a confusion matrix representing high true positive rates across all sessions. These outcomes asserted the effectiveness of semantic and structure aware techniques in discovering varied procedures of code plagiarism and highlighted the significance of incorporating graph theoretic methods with machine learning for robust taxonomy. The verdicts advocate for wider acceptance of semantic detection systems in educational technology, software forensics, and automated code review platforms.},
      keywords = {Code Plagiarism Detection, Abstract Syntax Tree (AST), Semantic Analysis, Graph Similarity, Machine Learning, Random Forest, Tree Edit Distance, Code Obfuscation, Multi-class Classification, Software Forensics},
      month = {July},
  }