Home / Current Issue / Paper 1713158
An AI-Based Smart Legal Assistant for Automated Legal Document Analysis
Subject area: Science,Engineering and Technology · Area of research: Artificial Intelligence and Machine Learning
Abstract
Understanding legal paperworks are a challenge as it's filled with hard-to-understand legal jargon and long sentences. Going through these papers by hand takes a lot of time and can lead to mixed results. Within this document is introduced the Smart Legal Assistant a system rooted in AI geared for the automation of legal document scrutiny through the avenues of extracting clauses, sorting them semantically, and evaluating risks. In practice the assistant is capable of taking both PDF and DOCX files breaking them down into segmented clauses by abiding to a set of predetermined rules and then goes on to classify each clause utilizing derivations from the models from transformer-based natural language processing techniques. Clauses are given a risk severity rating to point out terms that could be harmful or critical. FastAPI used in the backend for development and use of React in the frontend to make a review interface that interact with users. Exploring the output of the suggested setup it was observed that it notably slashes the time spent on going through documents and steps up the quality of how clear and consistent they are structured; all of these together these characteristics have made it a fitting choice for both academic work and the initial stages of legal reviews.
Keywords
Legal Document Analysis, Clause Classification, Transformer Models, Risk Assessment, Contract Review
References
[1] Chalkidis, Ilias, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. "LEGAL-BERT: The muppets straight out of law school." arXiv preprint arXiv:2010.02559 (2020).
[2] Hendrycks, Dan, Collin Burns, Anya Chen, and Spencer Ball. "Cuad: An expert-annotated nlp dataset for legal contract review." arXiv preprint arXiv:2103.06268 (2021)..
[3] Chalkidis, Ilias, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, and Nikolaos Aletras. "LexGLUE: A benchmark dataset for legal language understanding in English." In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 4310-4330. 2022.
[4] Koreeda, Yuta, and Christopher D. Manning. "ContractNLI: A dataset for document-level natural language inference for contracts." arXiv preprint arXiv:2110.01799 (2021).
[5] Zhang, Jingqing, Yao Zhao, Mohammad Saleh, and Peter Liu. "Pegasus: Pre-training with extracted gap-sentences for abstractive summarization." In International conference on machine learning, pp. 11328-11339. PMLR, 2020.
[6] Reimers, Nils, and Iryna Gurevych. "Sentence-bert: Sentence embeddings using siamese bert-networks." arXiv preprint arXiv:1908.10084 (2019).
[7] Xu, Y., M. Li, L. Cui, S. Huang, F. Wei, and M. Zhou. "LayoutLM: pre-training of text and layout for document image understanding. ArXiv." arXiv preprint arXiv:1912.13318 (2019).
[8] Xu, Yang, Yiheng Xuu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu et al. "Layoutlmv2: Multi-modal pre-training for visually-rich document understanding." In Proceeding of the 59th Annual Meeting of the Association for Computational Linguistic and the International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 2579-2591. 2021.
[9] Stanisławek, Tomasz, Filip Graliński, Anna Wróblewska, Dawid Lipiński, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemysław Biecek. "Kleister: key information extraction datasets involving long documents with complex layouts." In International Conference on Document Analysis and Recognition, pp. 564-579. Cham: Springer International Publishing, 2021.
[10] Garncarek, Łukasz, Rafał Powalski, Tomasz Stanisławek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Graliński. "Lambert: Layout-aware language modeling for information extraction." In International conference on document analysis and recognition, pp. 532-547. Cham: Springer International Publishing, 2021.
[11] Beltagy, Iz, Matthew E. Peters, and Arman Cohan. "Longformer: The long-document transformer."arXivpreprint arXiv:2004.05150 (2020).
[12] Lundberg, Scott M., and Su-In Lee. "A unified approach to interpreting model predictions." Advances in neural information processing systems 30 (2017).
[13] Zaheer, Manzil, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham et al. "Big bird: Transformers for longer sequences." Advances in neural information processing systems 33 (2020): 17283-17297.
[14] Mathew, Minesh, Ruben Tito, Dimosthenis Karatzas, R. Manmatha, and C. V. Jawahar. "Document visual question answering challenge 2020." arXiv preprint arXiv:2008.08899 (2020).
[15] Krpukhin, Vladimir, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. "Dense Passage Retrieval for Open-Domain Question Answering." In EMNLP (1), pp. 6769-6781. 2020.
How to cite this paper
@article{1713158,
author = {Rummana Firdaus , Manvitha P , Sunidhi S Babu , Sneha Nayak M S },
title = {An AI-Based Smart Legal Assistant for Automated Legal Document Analysis},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {6},
pages = {2090-2096},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1713158.pdf},
abstract = {Understanding legal paperworks are a challenge as it's filled with hard-to-understand legal jargon and long sentences. Going through these papers by hand takes a lot of time and can lead to mixed results. Within this document is introduced the Smart Legal Assistant a system rooted in AI geared for the automation of legal document scrutiny through the avenues of extracting clauses, sorting them semantically, and evaluating risks. In practice the assistant is capable of taking both PDF and DOCX files breaking them down into segmented clauses by abiding to a set of predetermined rules and then goes on to classify each clause utilizing derivations from the models from transformer-based natural language processing techniques. Clauses are given a risk severity rating to point out terms that could be harmful or critical. FastAPI used in the backend for development and use of React in the frontend to make a review interface that interact with users. Exploring the output of the suggested setup it was observed that it notably slashes the time spent on going through documents and steps up the quality of how clear and consistent they are structured; all of these together these characteristics have made it a fitting choice for both academic work and the initial stages of legal reviews.},
keywords = {Legal Document Analysis, Clause Classification, Transformer Models, Risk Assessment, Contract Review},
month = {December},
doi = {https://doi.org/10.64388/IREV9I6-1713158}
}