Home / Current Issue / Paper 1703529
Application of Python Programming Language in PDF-TEXT Based Information Extraction
Subject area: Science,Engineering and Technology · Area of research: Programming, Data Extraction
Abstract
Information Extraction has become a vital aspect of research over the decades which allow millions of researchers have access to only what seems important to them, amidst the enormous pieces of information around. The motivation of carrying this work out was to mitigate on the time factor that faces every researcher with regards to meeting stipulated deadlines. After a careful literature review on the existing studies it was gathered that a proposed system framework was developed to match keywords during the word extraction formation. The research followed a System Structured Analysis Design Methodology (SSADM), which utilizes the powerful libraries of python programming to mine text data known as keywords from a structured file and parse through these binary data. The extracted input data are encrypted via parallel encryption. The system was tested and the sample results achieved.
Keywords
Information, Extraction, Text data, Binary data, Encryption, Python
References
[1] Xiaofeng L., Zhiming Z.,(2019). Unsupervised Approaches for Textual Semantic Annotation, A Survey. https://dl.acm.org/doi/pdf/10.1145/3324473 University of Amsterdam=
[2] Gupta, V., Lehal, G.S.(2009): A survey of text mining techniques and applications. J. Emerg. Technol. Web Intell. 1(1), 60–
[3] Navathe, S.B., Ramez, E (2000).: Data warehousing and data mining. Fundam. Database Syst., 841– 872
[4] Said A. Salloum, Mostafa Al-Emran, Azza Abdel Monem and Khaled Shaalan(2016) Using Text Mining Techniques for Extracting Information from Research Articles Using Text Mining Techniques for Extracting Information from Research Articles Chapter in Studies in Computational Intelligence · January 2018 DOI: 10.1007/978-3-319-67056-0_18
[5] Zhang, Y., Chen, M., Liu, L.(2015): A review on text mining. 6th IEEE International Conference on Software Engineering and Service Science (ICSESS), pp. 681–685.
[6] M.Geetha, R C Pooja, J. Swetha, N. Nivedha, T. Daniya (2020) Implementation of Text Recognition and Text Extraction on Formatted Bills using Deep Learning International Journal of Control and Automation Vol. 13, No. 2, pp. 646 – 651
[7] T. Gnana Prakash K. Anusha (2017) Text Extraction from Image using Python International Journal of Trend in Scientific Research and Development (IJTSRD) International Open Access Journal, ISSN 2456-6470
[8] Xiaofeng L., Zhiming Z.,(2019). Unsupervised Approaches for Textual Semantic Annotation, A Survey. https://dl.acm.org/doi/pdf/10.1145/3324473 University of Amsterdam
[9] Sonit S., (2018). Natural Language Processing for Information Extraction. Department of Computing, Faculty of Science and Engineering, Macquarie University, Australia. arXiv:1807.02383v1 [cs.CL]
[10] Chengyao L., Huihua L., Yuanxing D., Yunliang C., (2016). Corpus based part-of-speech tagging. International Journal of speech Technology.https://www.readcube.com/articles/10.1007/s10772-016-9356-2
[11] Kong F., Xiaoming Z., (2010). Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). Association for Computational Linguistics. https://aclanthology.org/2021.acl-short.pdf
[12] Ly Papa A.,; Pedrinaci, C., Domingue J., (2012). Automated information extraction from web APIs documentation. The Open University’s repository of research publications and other research outputs. http://oro.open.ac.uk/34934/1/webApiDocProcessing.pdf
[13] Adnan., Akbar., (2018). An analytical study of information extraction from unstructured and multidimensional big data. Journal of big data. https://journalofbigdata.springeropen.com/articles/10.1186/s40537-019-0254-8
[14] Information Extraction (2020). Information ExtractionAbstracthttps://en.wikipedia.org/wiki/Information_extraction
How to cite this paper
@article{1703529,
author = {Benita A. Chinemerem, Donatus. O. Njoku, Taiwo Ahmed O.},
title = {Application of Python Programming Language in PDF-TEXT Based Information Extraction},
journal = {Iconic Research And Engineering Journals},
year = {2022},
volume = {6},
number = {3},
pages = {160-167},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1703529.pdf},
abstract = {Information Extraction has become a vital aspect of research over the decades which allow millions of researchers have access to only what seems important to them, amidst the enormous pieces of information around. The motivation of carrying this work out was to mitigate on the time factor that faces every researcher with regards to meeting stipulated deadlines. After a careful literature review on the existing studies it was gathered that a proposed system framework was developed to match keywords during the word extraction formation. The research followed a System Structured Analysis Design Methodology (SSADM), which utilizes the powerful libraries of python programming to mine text data known as keywords from a structured file and parse through these binary data. The extracted input data are encrypted via parallel encryption. The system was tested and the sample results achieved.},
keywords = {Information, Extraction, Text data, Binary data, Encryption, Python},
month = {September},
}