Home / Current Issue / Paper 1718010
HackEval: An Intelligent Multi-Agent Framework for Automated, Bias- Mitigated Assessment in Competitive Hackathon Ecosystems
Subject area: Science,Engineering and Technology · Area of research: Artificial Intelligence and Multi-Agent Systems
DOI: 10.64388/IREV9I11-1718010
Abstract
Contemporary hackathon adjudication is burdened by four structural deficiencies inherent to human- centric evaluation: inconsistent rubric application, substantial inter-rater score variance, prohibitive assessment latency, and a near-total absence of granular, actionable post-event diagnostic feedback. This paper introduces HackEval, a production-grade multi-agent artificial intelligence framework designed to systematica ly resolve these limitations through real-time, bias-mitigated evaluation across heterogeneous project submission modalities. Six functionally specialized agents operate in parallel: (i) a Code Quality Agent performing deep multi- criterion static analysis on GitHub repositories; (ii) a Presentation Analyzer Agent realizing a four-stage pipeline integrating LLM semantic reasoning with contrastive vision-language embeddings; (iii) a UI/UX Evaluation Agent leveraging CLIP-based aesthetic regression [13]; (iv) an Innovation Agent quantifying originality via semantic embedding distance; (v) a Fea sibility Agent applying chain-of-thought LLM reasoning [12]; and (vi) a Plagiarism Detection Agent employing sentence- transformer cosine similarity with FAISS indexing [15]. Empirical evaluation across 27 authentic hackathon submissions yields a Pearson correlation of r = 0.93 between AI composite scores and consensus expert evaluations, a 92.8% reduction in per-team evaluation time, and a 77.4% improvement in cross-evaluator scoring consistency. The platform is delivered as a multi-tenant Software-as-a-Service system built on a MERN + FastAPI + LangChain stack, demonstrating concurrent scalability exceeding 500 simultaneous teams with sub-10-second feedback delivery.
Keywords
Automated Code Assessment, Bias Reduction Mechanisms; Competitive Hackathon Evaluation; Large Language Model Scoring; Multi-Agent AI Systems; Real-Time Feedback Systems; Saa S Architecture; Semantic Embeddings; Static Code Analysis; Vision-Language Models.
References
[1] HackerEarth Inc., “Global hackathon trends report 2023,” HackerEarth, San Francisco, CA, USA, Tech. Rep., 2023. [Online]. Available: https://www.hackerearth.com/hackathon-report
[2] R. Stemler and J. Tsai, “Measuring inter-rater reliability in competition judging: A meta-analysis,” J. Educ. Meas., vol. 41, no. 2, pp. 113–128, 2020, doi: 10.1111/jedm.12254.
[3] Devpost Inc., “Devpost — The home for hackathons,” 2024. [Online]. Available: https://devpost.com/
[4] HackerEarth Inc., “HackerEarth hackathons,” 2024. [Online].Available:
[5] Unstop (Dare2Compete), “Unstop competitions platform,” 2024. [Online]. Available: https://unstop.com/
[6] Kaggle Inc., “Kaggle competitions,” 2024. [Online]. Available: https://www.kaggle.com/competitions/
[7] P. Ihantola, T. Ahoniemi, V. Karavirta, and O. Seppälä, “Review of recent systems for automatic assessment of programming assignments,” in Proc. 10th Koli Calling Int. Conf. Comput. Educ. Res., Koli, Finland, 2010, pp. 86–93.
[8] Y. Liang et al., “Can large language models write good code? An empirical study,” in Proc. ACM SIGSOFT Int. Symp. Softw. Testing Anal. (ISSTA), Seattle, WA, USA, 2023.
[9] K. Stasaski and M. Hearst, “Automatically scoring explanations in science education,” in Findings Assoc. Comput. Linguistics: EMNLP 2022, Abu Dhabi, UAE, 2022, pp. 4463–4476.
[10] A. V. Savchenko, “Rapid attributes-based image retrieval using fine-tuned vision-language models,” Neural Netw., vol. 169, pp. 222–234, 2024.
[11] C. Mao, Y. Chen, and T. Liu, “Evaluating novelty and originality in generated text using embedding distance metrics,” in Proc. AAAI Workshop AI Eval., Washington, DC, USA, 2023.
[12] T. Brown et al., “Language models are few-shot learners,” in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, pp. 1877– 1901, 2020.
[13] A. Radford et al., “Learning transferable visual models from natural language supervision,” in Proc. 38th Int. Conf. Mach. Learn. (ICML), 2021, pp. 8748–8763.
[14] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using siamese BERT-networks,” in Proc. 2019 Conf. Empirical Methods Natural Language Process. (EMNLP), Hong Kong, China, 2019, pp. 3982–3992.
[15] J. Johnson, M. Douze, and H. Jégou, “Billion-scale similarity search with GPUs,” IEEE Trans. Big Data, vol. 7, no. 3, pp. 535–547, 2021.
How to cite this paper
@article{1718010,
author = {Rishabh Tripathi, Sanskrati Agrawal},
title = {HackEval: An Intelligent Multi-Agent Framework for Automated, Bias- Mitigated Assessment in Competitive Hackathon Ecosystems},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {11},
pages = {3338-3346},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1718010.pdf},
abstract = {Contemporary hackathon adjudication is burdened by four structural deficiencies inherent to human- centric evaluation: inconsistent rubric application, substantial inter-rater score variance, prohibitive assessment latency, and a near-total absence of granular, actionable post-event diagnostic feedback. This paper introduces HackEval, a production-grade multi-agent artificial intelligence framework designed to systematica ly resolve these limitations through real-time, bias-mitigated evaluation across heterogeneous project submission modalities. Six functionally specialized agents operate in parallel: (i) a Code Quality Agent performing deep multi- criterion static analysis on GitHub repositories; (ii) a Presentation Analyzer Agent realizing a four-stage pipeline integrating LLM semantic reasoning with contrastive vision-language embeddings; (iii) a UI/UX Evaluation Agent leveraging CLIP-based aesthetic regression [13]; (iv) an Innovation Agent quantifying originality via semantic embedding distance; (v) a Fea sibility Agent applying chain-of-thought LLM reasoning [12]; and (vi) a Plagiarism Detection Agent employing sentence- transformer cosine similarity with FAISS indexing [15]. Empirical evaluation across 27 authentic hackathon submissions yields a Pearson correlation of r = 0.93 between AI composite scores and consensus expert evaluations, a 92.8% reduction in per-team evaluation time, and a 77.4% improvement in cross-evaluator scoring consistency. The platform is delivered as a multi-tenant Software-as-a-Service system built on a MERN + FastAPI + LangChain stack, demonstrating concurrent scalability exceeding 500 simultaneous teams with sub-10-second feedback delivery.},
keywords = {Automated Code Assessment, Bias Reduction Mechanisms; Competitive Hackathon Evaluation; Large Language Model Scoring; Multi-Agent AI Systems; Real-Time Feedback Systems; Saa S Architecture; Semantic Embeddings; Static Code Analysis; Vision-Language Models.},
month = {May},
doi = {https://doi.org/10.64388/IREV9I11-1718010}
}