Home / Current Issue / Paper 1706049
Developing & Comparing Various Topic Modeling Algorithms on a Stack Overflow Dataset
Subject area: Science,Engineering and Technology · Area of research: Topic Modelling
Abstract
This research extracts and compares coherent and instructive topics from Stack Overflow, a tech and programming community. Scraping questions, summaries, and tags, exploratory data analysis, rigorous pre-processing, and topic models to find latent topics are the study's main steps. LSA, LDA, and BERTopic are popular topic models. To achieve the best models for each algorithm, base model hyperparameters were tweaked and refined. Then, each algorithm's models were compared for performance and accuracy using coherence score, topic distinctiveness, and different visualization techniques to examine semantic separation. Each technique was tested to see how well it handled different data dimensions. The comparison study showed that BERTopic was the best topic model, achieving more granular and semantically meaningful categorizations through improved semantic comprehension, topic distinguishability, and topic extraction coherence. This research shows how advanced topic modelling may extract nuanced insights from text data, giving a complete process from data acquisition to subject categorization. The results demonstrate BERTopic's ability to decipher complicated textual relationships and generate coherent words for varied themes. Thus, this research improves information retrieval and user experience on online community platforms like Stack Overflow by using advanced natural language processing models.
Keywords
BERTopic, Latent Semantic Allocation, Latent Dirichlet Allocation
References
[1] Alghamdi, R. and Alfalqi, K. (2015) ‘A Survey of Topic Modeling in Text Mining’, International Journal of Advanced Computer Science and Applications, 6(1), pp. 147–153.
[2] Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet Allocation. Journal of Machine Learning Research 3, 3(2), 993-1022.
[3] David, C. (2023, January 17). Stack Overflow Growth and Usage Statistics (2023). Retrieved from usesignhouse: https://www.usesignhouse.com/blog/stack-overflow-stats
[4] Deerwester, S., Dumais, S. T., Furnas, G. W., Landauer, T. K., & Harshman, R. (1990, Sep). Indexing by Latent Semantic Analysis. Journal for the American Society for Information Science, 41(6).
[5] Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. North American Chapter for the Association of Computational Linguistics
[6] Hofmann, T. (1999). Probabilistic latent semantic indexing. SIGIR '99: Proceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrieval, (pp. 50–57).
[7] Kherwa, P. and Bansal, P. (2020) ‘Topic Modeling: A Comprehensive Review’, EAI Transactions on Scalable Information Systems, 7(24), pp. 1–16.
[8] Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space.
[9] Salton, G. (1986). Recent trends in automatic information retrieval. Wikisym05: Int'l Symposium on Wikis Palazzo dei Congressi. Pisa Italy: Association for Computing Machinery New York, United States.
[10] Pennington, J., Socher, R., & Manning, C. D. (2014). GloVe: Global Vectors for Word Representation. Conference on Empirical Methods in Natural Language Processing (EMNLP), (pp. 1532–1543)
[11] UVWghj|}’“”•›œßàáâêôåôåØÍ÷ªÃ·ªÃ·ÍœŽœ‚vj[I"he¯h^~ 5�6�CJOJQJaJh^~ h˜gCJOJQJaJh^~ CJOJQJaJhe¯h^~ 6�OJQJhe¯hØ}N6�OJQJhe¯he¯6�H*OJQJhe¯h^~ 6�H*OJQJh^~ he¯OJPJQJh^~ he¯H*OJQJhe¯OJPJQJh^~ h^~ OJQJh^~ h^~ CJOJQJhØ}NhØ}NCJ(OJPJQJhØ}NCJ(OJPJQJVW”àáâwxÉËØÙ¯°‘õèÓõÁÁ···· ·······$ &F„7„Éý¤^„7`„Éýa$gde¯m$ $¤a$gde¯„ñÿ„ dð¤]„ñÿ^„ gd^~ $„`„öÿdë¤^„``„öÿa$gd˜g$dð¤a$gd^~ $¤a$gde¯êìó ª«« [Some characters in this reference could not be displayed correctly — please refer to the published PDF for the full reference.]
[12] vwx}…‡ÈÉÊËרÙìÔÔ¿ÔÔÔ«ì«ì™‡xfTDD3!he¯CJOJQJ\�aJmH sH he¯h^~ CJOJQJ\�aJ"he¯h^~ 6�CJOJQJ\�aJ"he¯hÄ|n6�CJOJQJ\�aJhe¯6�CJOJQJ\�aJ"he¯hÄ|n5�6�CJOJQJaJ"he¯h^~ 5�6�CJOJQJaJ&he¯he¯5�6�CJOJPJQJaJ(he¯5�6�CJOJPJQJaJmH sH .he¯hÄ|n5�6�CJOJPJQJaJmH sH &he¯h^~ 5�6�CJOJPJQJaJ�‘jkª;$ku}‹kx[\]nopëÚëÆëëÚëëëëÚëÚëëëëë²ëëëë¡�}p`he¯he¯CJOJQJ\�aJh2dXCJOJQJ\�aJhe¯h2dXCJOJQJ\�aJ'he¯he¯CJOJQJ\�aJmH sH !h¸OCJOJQJ\�aJmH sH 'he¯h<JåCJOJQJ\�aJmH sH 'he¯hcu¹CJOJQJ\�aJmH sH !he¯CJOJQJ\�aJmH sH 'he¯h¸OCJOJQJ\�aJmH sH ‘jkª;$‹x\]op§¨opR S Ì!Í!ß#à#ú&"'P'õõõõõõõõõÞÓõõõõõõõõõõõõõhe¯h!|âCJOJQJ\�aJ+he¯h2dXCJOJPJQJ\�aJnH tH he¯hß1¸CJOJQJ\�aJhe¯CJOJQJ\�aJhe¯h#3ŠCJOJQJ\�aJhe¯h‚ÏCJOJQJ\�aJhe¯hcu¹CJOJQJ\�aJhe¯h2dXCJOJQJ\�aJ*he¯hcu¹CJOJQJ\�aJmHnHu$&he¯hN~gCJOJQJ\�aJ"he¯hN~gCJOJQJ\�]�aJ,jèhe¯hß1¸CJEHìÿOJQJU\�aJ,jhe¯hß1¸CJEHìÿOJQJU\�aJ"he¯hß1¸CJOJQJ\�]�aJ+jhe¯hß1¸CJOJQJU\�]�aJ+he¯hß1¸CJOJPJQJ\�aJnH tH he¯hß1¸CJOJQJ\�aJ%he¯hß1¸6�CJOJQJ\�]�aJV'W'�'’'“'”'Å'Æ'à'›(ž(Ÿ( (¡(¥(ª(&*'*J*O*',C,F,H,I,{,�,‡-�-ì-í-û-4.6.h.j.l.n.~.€.¸.º./ /L/N/P/R/T/X/�/–/˜/š/¨/ª/¸/ìÜìÜìÜìÜÜÌÌÜ¿ÜìÜ¿ÜìÜÜÌÜ¿ÜìÜìܲÜÜìÜìÜìÜìÜìÜìÜìÜìÜìŸ�ܲÜìÜhe¯h+qDCJOJQJ\�aJ%he¯h+qD6�CJOJQJ\�]�aJh$q'CJOJQJ\�aJhe¯CJOJQJ\�aJhe¯h!|âCJOJQJ\�aJhe¯hN~gCJOJQJ\�aJ%he¯hN~g6�CJOJQJ\�]�aJ8P' (¡(&*'*H,I,ì-í-L/˜/š/¨/~0¢192:2Ä2Å2…3†3o8p8ÿ<=.?õõõõõõõõõõõõÜÜÜõõõõõõõõõõ$ &F [Some characters in this reference could not be displayed correctly — please refer to the published PDF for the full reference.]
How to cite this paper
@article{1706049,
author = {Raphael Ibraimoh, Kwame Ofosu Debrah, Emmanuel Nwambuowo},
title = {Developing & Comparing Various Topic Modeling Algorithms on a Stack Overflow Dataset},
journal = {Iconic Research And Engineering Journals},
year = {2024},
volume = {8},
number = {1},
pages = {243-253},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1706049.pdf},
abstract = {This research extracts and compares coherent and instructive topics from Stack Overflow, a tech and programming community. Scraping questions, summaries, and tags, exploratory data analysis, rigorous pre-processing, and topic models to find latent topics are the study's main steps. LSA, LDA, and BERTopic are popular topic models. To achieve the best models for each algorithm, base model hyperparameters were tweaked and refined. Then, each algorithm's models were compared for performance and accuracy using coherence score, topic distinctiveness, and different visualization techniques to examine semantic separation. Each technique was tested to see how well it handled different data dimensions. The comparison study showed that BERTopic was the best topic model, achieving more granular and semantically meaningful categorizations through improved semantic comprehension, topic distinguishability, and topic extraction coherence. This research shows how advanced topic modelling may extract nuanced insights from text data, giving a complete process from data acquisition to subject categorization. The results demonstrate BERTopic's ability to decipher complicated textual relationships and generate coherent words for varied themes. Thus, this research improves information retrieval and user experience on online community platforms like Stack Overflow by using advanced natural language processing models.},
keywords = {BERTopic, Latent Semantic Allocation, Latent Dirichlet Allocation},
month = {July},
}