International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1703529

1703529 Vol 6 · Issue 3 Download Paper

Application of Python Programming Language in PDF-TEXT Based Information Extraction

Benita A. Chinemerem Donatus. O. Njoku Taiwo Ahmed O.

Subject area: Science,Engineering and Technology  ·  Area of research: Programming, Data Extraction

Abstract

Information Extraction has become a vital aspect of research over the decades which allow millions of researchers have access to only what seems important to them, amidst the enormous pieces of information around. The motivation of carrying this work out was to mitigate on the time factor that faces every researcher with regards to meeting stipulated deadlines. After a careful literature review on the existing studies it was gathered that a proposed system framework was developed to match keywords during the word extraction formation. The research followed a System Structured Analysis Design Methodology (SSADM), which utilizes the powerful libraries of python programming to mine text data known as keywords from a structured file and parse through these binary data. The extracted input data are encrypted via parallel encryption. The system was tested and the sample results achieved.

Keywords

Information, Extraction, Text data, Binary data, Encryption, Python

References

[1] Xiaofeng L., Zhiming Z.,(2019). Unsupervised Approaches for Textual Semantic Annotation, A Survey. https://dl.acm.org/doi/pdf/10.1145/3324473 University of Amsterdam=

[2] Gupta, V., Lehal, G.S.(2009): A survey of text mining techniques and applications. J. Emerg. Technol. Web Intell. 1(1), 60–

[3] Navathe, S.B., Ramez, E (2000).: Data warehousing and data mining. Fundam. Database Syst., 841– 872

[4] Said A. Salloum, Mostafa Al-Emran, Azza Abdel Monem and Khaled Shaalan(2016) Using Text Mining Techniques for Extracting Information from Research Articles Using Text Mining Techniques for Extracting Information from Research Articles Chapter  in  Studies in Computational Intelligence · January 2018 DOI: 10.1007/978-3-319-67056-0_18

[5] Zhang, Y., Chen, M., Liu, L.(2015): A review on text mining. 6th IEEE International Conference on Software Engineering and Service Science (ICSESS), pp. 681–685.

[6] M.Geetha, R C Pooja, J. Swetha, N. Nivedha, T. Daniya (2020) Implementation of Text Recognition and Text Extraction on Formatted Bills using Deep Learning International Journal of Control and Automation Vol. 13, No. 2, pp. 646 – 651

[7] T. Gnana Prakash K. Anusha (2017) Text Extraction from Image using Python International Journal of Trend in Scientific Research and Development (IJTSRD) International Open Access Journal, ISSN 2456-6470

[8] Xiaofeng L., Zhiming Z.,(2019). Unsupervised Approaches for Textual Semantic Annotation, A Survey. https://dl.acm.org/doi/pdf/10.1145/3324473 University of Amsterdam

[9] Sonit S., (2018). Natural Language Processing for Information Extraction. Department of Computing, Faculty of Science and Engineering, Macquarie University, Australia. arXiv:1807.02383v1 [cs.CL]

[10] Chengyao L., Huihua L., Yuanxing D., Yunliang C., (2016). Corpus based part-of-speech tagging. International Journal of speech Technology.https://www.readcube.com/articles/10.1007/s10772-016-9356-2

[11] Kong F., Xiaoming Z., (2010). Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). Association for Computational Linguistics. https://aclanthology.org/2021.acl-short.pdf

[12] Ly Papa A.,; Pedrinaci, C., Domingue J., (2012). Automated information extraction from web APIs documentation. The Open University’s repository of research publications and other research outputs. http://oro.open.ac.uk/34934/1/webApiDocProcessing.pdf

[13] Adnan., Akbar., (2018). An analytical study of information extraction from unstructured and multidimensional big data. Journal of big data. https://journalofbigdata.springeropen.com/articles/10.1186/s40537-019-0254-8

[14] Information Extraction (2020). Information ExtractionAbstracthttps://en.wikipedia.org/wiki/Information_extraction

How to cite this paper

Benita A. Chinemerem, Donatus. O. Njoku, Taiwo Ahmed O. "Application of Python Programming Language in PDF-TEXT Based Information Extraction" Iconic Research And Engineering Journals Volume 6 Issue 3 2022 Page 160-167
Benita A. Chinemerem, Donatus. O. Njoku, Taiwo Ahmed O. "Application of Python Programming Language in PDF-TEXT Based Information Extraction" Iconic Research And Engineering Journals, vol. 6, no. 3, Sep. 2022
Benita A. Chinemerem, Donatus. O. Njoku, Taiwo Ahmed O. (2022). Application of Python Programming Language in PDF-TEXT Based Information Extraction. Iconic Research And Engineering Journals, 6(3).
Benita A. Chinemerem, Donatus. O. Njoku, Taiwo Ahmed O. "Application of Python Programming Language in PDF-TEXT Based Information Extraction" Iconic Research And Engineering Journals, vol. 6, no. 3, Sep. 2022.
@article{1703529,
      author = {Benita A. Chinemerem, Donatus. O. Njoku, Taiwo Ahmed O.},
      title = {Application of Python Programming Language in PDF-TEXT Based Information Extraction},
      journal = {Iconic Research And Engineering Journals},
      year = {2022},
      volume = {6},
      number = {3},
      pages = {160-167},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1703529.pdf},
      abstract = {Information Extraction has become a vital aspect of research over the decades which allow millions of researchers have access to only what seems important to them, amidst the enormous pieces of information around. The motivation of carrying this work out was to mitigate on the time factor that faces every researcher with regards to meeting stipulated deadlines. After a careful literature review on the existing studies it was gathered that a proposed system framework was developed to match keywords during the word extraction formation. The research followed a System Structured Analysis Design Methodology (SSADM), which utilizes the powerful libraries of python programming to mine text data known as keywords from a structured file and parse through these binary data. The extracted input data are encrypted via parallel encryption. The system was tested and the sample results achieved.},
      keywords = {Information, Extraction, Text data, Binary data, Encryption, Python},
      month = {September},
  }