Home / Current Issue / Paper 1712709
DataScribe: An Automated EDA and Narrative Reporting Framework for Accessible Data Analysis
Subject area: Science,Engineering and Technology · Area of research: Data Science & Human-Computer Interaction
Abstract
Exploratory Data Analysis (EDA) remains time-intensive and inaccessible to non-technical users despite its critical role in data science workflows. DataScribe addresses this gap through an automated pipeline that generates visualizations, statistical summaries, and human-readable narrative explanations from uploaded CSV/Excel datasets. The system produces multi-format reports (PDF, HTML, Excel, R-code) in under 12 seconds. Testing on the Titanic dataset (891 rows, 12 columns) demonstrated 83% reduction in analysis time compared to manual approaches, with 87% of non-technical users successfully interpreting results without statistical training. Deployed at https://datascribe.onrender.com/, the system bridges the accessibility gap in data analysis through automated narrative generation and reproducible code export.
Keywords
Automated EDA, Data Visualization, Narrative Reporting, Python-R Integration, Data Storytelling
References
[1] Islam, S., et al. (2024). DataNarrative: Automated Data-Driven Storytelling with Visualizations and LLMs. Proc. ACM CHI, 45(3), 234-248.
[2] Manatkar, A., et al. (2024). QUIS: Question-Guided Insights for Automated EDA. J. Data Sci. Analytics, 12(2), 145-162.
[3] Wongsuphasawat, K., et al. (2016). Voyager: Exploratory Analysis via Faceted Browsing. IEEE TVCG, 22(1), 649-658.
[4] Vartak, M., et al. (2017). Towards a System for Automatic EDA. Proc. Workshop HILDA, Article 5.
[5] Dibia, V., & Demiralp, C. (2019). Data2Vis: Automatic Visualization Generation. IEEE CG&A, 39(5), 33-46.
[6] Kandel, S., et al. (2012). Enterprise Data Analysis: An Interview Study. IEEE TVCG, 18(12), 2917-2926.
[7] Satyanarayan, A., et al. (2017). Vega-Lite: A Grammar of Interactive Graphics. IEEE TVCG, 23(1), 341-350.
[8] Tukey, J. W. (1977). Exploratory Data Analysis. Addison-Wesley.
How to cite this paper
@article{1712709,
author = {Anushree, Sakun Choudhary, Mansi Lakhmani, Roopali Gupta},
title = {DataScribe: An Automated EDA and Narrative Reporting Framework for Accessible Data Analysis},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {6},
pages = {892-896},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1712709.pdf},
abstract = {Exploratory Data Analysis (EDA) remains time-intensive and inaccessible to non-technical users despite its critical role in data science workflows. DataScribe addresses this gap through an automated pipeline that generates visualizations, statistical summaries, and human-readable narrative explanations from uploaded CSV/Excel datasets. The system produces multi-format reports (PDF, HTML, Excel, R-code) in under 12 seconds. Testing on the Titanic dataset (891 rows, 12 columns) demonstrated 83% reduction in analysis time compared to manual approaches, with 87% of non-technical users successfully interpreting results without statistical training. Deployed at https://datascribe.onrender.com/, the system bridges the accessibility gap in data analysis through automated narrative generation and reproducible code export.},
keywords = {Automated EDA, Data Visualization, Narrative Reporting, Python-R Integration, Data Storytelling},
month = {December},
doi = {https://doi.org/10.64388/IREV9I6-1712709}
}