Article Detail

Loading...

Developing a Web Platform to Extract Information from Receipt Images

Barış DEMİRHAN*, Yakup KUTLU

Keywords

OCR OpenCV Digitization LLM Regex Receipt Web platform

Doi : 10.71350/jner.2026166

Abstract

Data extracted from receipts are necessary for various applications in many industries, and this extraction process can be simplified through pre-trained deep learning models. However, high-quality receipt images are essential for algorithms to produce accurate models. This paper focuses on the techniques to improve the image preprocessing process and the information extraction methods applied in Turkish and English and create a web-based receipt scanning and data management system. Optical Character Recognition (OCR) technology is a system that provides a full alphanumeric recognition of printed or handwritten characters from images. Initially, OpenCV has been used to detect the bill or invoice and filter out the unnecessary noise from the image. Then the intermediate image is passed for further processing using the Tesseract OCR engine, which is an optical character recognition engine. This application is built upon a Flask backend, a responsive HTML/CSS/JavaScript frontend, and a SQLite database, offering capabilities including single and batch receipt scanning, editable JSON-based data review by both using regex and AI parser, persistent storage across two language-specific database tables, Excel export, and downloadable output bundles including raw OCR text and structured JSON files. Our methodology and system prove to be highly accurate while tested on a variety of input images of bills and invoices. The architecture is modular and extensible, making it suitable as a foundation for larger-scale document digitization workflows.

References

  1. Ahmed, S., Ali, A., & Naser, E. (2023). Tesseract OpenCV versus CNN: A comparative study on the recognition of unified modern Iraqi license plates. Revue d’Intelligence Artificielle, 37(5), 727–734.
  2. Aydin, T., Yildizak, B., & Caliskan, E. E. (2024). A comparative study of layout based and OCR-free models for Turkish receipt documents. Proceedings of the 2024 IEEE Conference. https://doi.org/10.1109/2024.turkish.receipts
  3. Garcia, M. B., & Claour, J. P. (2021, November). Mobile bookkeeper: Personal financial management application with receipt scanner using optical character recognition. Proceedings of the 2021 1st Conference on Online Teaching for Mobile Education (OT4ME), 15–20.
  4. Kamisetty, V. N. S. R., Chidvilas, S., Revathy, S., Jeyanthi, P., Anu, V. M., & Gladence, L. M. (2022). Digitization of data from invoice using OCR. Proceedings of the 2022 6th International Conference on Computing Methodologies and Communication (ICCMC), 1–8. https://doi.org/10.1109/ICCMC53470.2022.9754117
  5. Kumar, V., Kaware, P., Singh, P., Sonkusare, R., & Kumar, S. (2020). Extraction of information from bill receipts using optical character recognition. Proceedings of the 2020 International Conference on Smart Electronics and Communication (ICOSEC), 72–77. https://doi.org/10.1109/ICOSEC49089.2020.9215246
  6. Nguyen, H. Q., Vo-Xuan, T., Nguyen, C. T., Ngo, T., Nguyen, D. T. V., & Huynh, K. T. (2024). Multilingual receipt image preprocessing optimization for OCR. Proceedings of the 2024 International Conference on Advanced Technologies for Communications (ATC), 922–927. https://doi.org/10.1109/ATC63255.2024.10908135
  7. Opencv.org. (2020). Information related to Open CV. https://opencv.org/
  8. Ozdil, M. A., & Vural, F. T. Y. (1997). Optical character recognition without segmentation. Proceedings of the Fourth International Conference on Document Analysis and Recognition, 2, 483–486. https://doi.org/10.1109/ICDAR.1997.620545
  9. Rabby, A. S. A., Islam, M. M., Hasan, N., Nahar, J., & Rahman, F. (2020, July). Language detection using convolutional neural network. Proceedings of the 2020 11th International Conference on Computing, Communication and Networking Technologies (ICCCNT), 1–5.
  10. Smith, R. (2007). An overview of the Tesseract OCR engine. Proceedings of the Ninth International Conference on Document Analysis and Recognition (ICDAR 2007), 629–633. https://doi.org/10.1109/ICDAR.2007.4376991

40 13

Article Summery

ISSN : 3108-6438

Volume 2 Issue 1

Submission Date: 2026-05-13

Accepted Date : 2026-06-23

Available Online : 2026-06-30

Publication Date :2026-06-30



How to Cite

Cite as :

DEMİRHAN, B., KUTLU, Y. (2026). Developing a Web Platform to Extract Information from Receipt Images. Journal of Natural and Engineering Research, 2(1), 56-71, doi : 10.71350/jner.2026166