OCR Table Extraction and Compliance Verification of Legal Documents Based on Multimodal Transformer
Main Article Content
Abstract
Accurate extraction and intelligent interpretation of structured information from complex documents remain challenging due to heterogeneous data modalities, intricate layouts, and long-range semantic dependencies. This study proposes a multimodal document understanding framework integrating optical character recognition (OCR), multimodal Transformer encoding, graph neural networks (GNNs), and reinforcement-learning-based reasoning. A dedicated OCR and layout-analysis module is first employed to acquire textual, visual, and spatial information from complex document images. Subsequently, a multimodal Transformer encoder performs joint representation learning by fusing visual features, semantic content, and layout embeddings to achieve robust table structure reconstruction and structured information extraction. To model semantic relationships among extracted entities, a relational graph convolutional network is constructed for heterogeneous information representation and propagation. Furthermore, a rule-guided reinforcement learning mechanism is introduced to optimize reasoning paths and improve decision accuracy in complex verification tasks. Experimental results demonstrate that the proposed framework achieves a GriTS score of 0.927 for table reconstruction and an F1 score of 93.1% for end-to-end verification, outperforming existing document-understanding approaches while significantly improving processing efficiency. The proposed framework provides an effective methodology for multimodal information fusion, graph-based information propagation, intelligent signal representation, and adaptive reasoning in complex information-processing systems, offering potential applications in intelligent sensing, communication-oriented information analysis, and distributed decision-support environments.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
S. Joseph, “Advanced digital image processing technique based optical character recognition of scanned document,” Journal of Innovative Image Processing, vol. 4, no. 3, pp. 195-205, 2022, doi: 10.36548/jiip.2022.3.007.
S. Mukherjee, H. Tyagi, P. Tyagi, et al., “OCR using python and its application,” Journal of Computer Science, vol. 12, no. 3, pp. 45-58, 2023, doi: 10.17762/jaz.v44iS-3.1062.
A. Chawla, A. Gupta, and S. Shushrutha K, “Intelligent information retrieval: techniques for character recognition and structured data extraction,” J. Emerg. Technol. Innov. Res.(JETIR), vol. 9, no. 7, pp. 452-459, 2022.
R. Samantapudi R K, “Table Extraction from Financial and Transactional Documents,” International journal of IoT, vol. 5, no. 01, pp. 95-125, 2025, doi: 10.55640/ijiot-05-01-06.
K. Pappula K and P. Rusum G, “Multi-Modal AI for Structured Data Extraction from Documents,” International Journal of Emerging Research in Engineering and Technology, vol. 4, no. 3, pp. 75-86, 2023, doi: 10.63282/3050-922X.IJERET-V4I3P109.
K. Pappula K, “Transformer-Based Classification of Financial Documents in Hybrid Workflows,” International Journal of Multidisciplinary on Science and Management, vol. 1, no. 3, pp. 48-61, 2024.
Z. Zhang, K. Chen, R. Wang, et al., “Universal multimodal representation for language understanding,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 7, pp. 9169-9185, 2023, doi: 10.1109/TPAMI.2023.3234170.
B. Oral and G. Eryiğit , “Fusion of visual representations for multimodal information extraction from unstructured transactional documents,” International Journal on Document Analysis and Recognition (IJDAR), vol. 25, no. 3, pp. 187-205, 2022, doi: 10.1007/s10032-022-00399-3.
L. Polo, “The role of AI and OCR-based label verification systems in enhancing food traceability and supply chain transparency,” International Journal for Multidisciplinary Research, vol. 7, no. 2, pp. 1-18, 2025, doi: 10.36948/ijfmr.2025.v07i02.41577.
S. Dutta, S. Adhikary, and D. Dwivedi A, “Visformers-Combining vision and transformers for enhanced complex document classification,” Machine Learning and Knowledge Extraction, vol. 6, no. 1, pp. 448-463, 2024, doi: 10.3390/make6010023.
Q. Zhang, “Enhanced Feature Fusion and Transfer Learning for Multi-Format Government Document Classification,” Journal of Science, Innovation & Social Impact, vol. 1, no. 1, pp. 427-441, 2025.
S. Maldini A, J. Saputra W S, and A. Prasetya D, “Multimodal Detection of Covert Online Gambling Advertisements Using Faster R-CNN and Tr-OCR,” bit-Tech, vol. 8, no. 1, pp. 953-963, 2025, doi: 10.32877/bt.v8i1.2769.
B. Bulatov K, V. Bezmaternykh P, P. Nikolaev D, et al., “Towards a unified framework for identity documents analysis and recognition,” Kompyuternaya optika, vol. 46, no. 3, pp. 436-454, 2022, doi: 10.18287/2412-6179-CO-1024.
S. Mandvikar, “Augmenting intelligent document processing (IDP) workflows with contemporary large language models (LLMs),” International Journal of Computer Trends and Technology, vol. 71, no. 10, pp. 80-91, 2023, doi: 10.14445/22312803/IJCTT-V71I10P110.
T. Odofin, A. Abayomi A, E. Ogbuefi, et al., “Strategic integration of LangChain, Hugging Face Transformers, and OpenAI for document intelligence systems,” International Journal of Scientific Research in Science Engineering and Technology, vol. 11, no. 4, pp. 413-426, 2024, doi: 10.32628/IJSRSET25121177.
V. Naik, P. Patel, and R. Kannan, “Legal entity extraction: An experimental study of NER approach for legal documents,” International Journal of Advanced Computer Science and Applications, vol. 14, no. 3, pp. 775-783, 2023, doi: 10.14569/IJACSA.2023.0140389.
B. Sinha, S. Saxena, A. Vashishtha, et al., “Legal Ease: An End-to-End Automated Legal Document Processing System Using LLMs and OCR,” International Journal of Software Computing and Testing, vol. 11, no. 1, pp. 29-32p, 2025.
J. Chen, M. Wang, and T. Sun, “Intelligent Tax Systems and the Role of Natural Language Processing in Regulatory Interpretation,” American Journal of Machine Learning, vol. 6, no. 4, pp. 74-94, 2025, doi: 10.71465/ajml3452.
T. Zhang, “A Knowledge Graph-Enhanced Multimodal AI Framework for Intelligent Tax Data Integration and Compliance Enhancement,” Frontiers in Business and Finance, vol. 2, no. 02, pp. 247-261, 2025, doi: 10.71465/fbf384.
V. Bezditnyi, “Use of artificial intelligence for tax planning optimization and regulatory compliance,” Research Corridor Journal of Engineering Science, vol. 1, no. 1, pp. 103-142, 2024, doi: 10.66320/8qe5f566.
A. Parakala and R. Pothula, “AI+ Document Understanding in UiPath: Solving Real Government Problems,” International Journal of Artificial Intelligence, Data Science, and Machine Learning, vol. 3, no. 3, pp. 111-122, 2022, doi: 10.63282/3050-9262.IJAIDSML-V3I3P112.
R. Pingili, “AI-driven intelligent document processing for banking and finance,” International Journal of Management & Entrepreneurship Research, vol. 7, no. 2, pp. 98-109, 2025, doi: 10.51594/ijmer.v7i2.1802.
S. Talakola, “Transforming BOL Images into Structured Data Using AI,” International Journal of Artificial Intelligence, Data Science, and Machine Learning, vol. 6, no. 1, pp. 105-114, 2025, doi: 10.63282/3050-9262.IJAIDSML-V6I1P112.
A. Chansarkar, “Enhancing Unstructured Document Text Extraction with LLMs in Cloud-Native Enterprise Solutions,” Journal Of Engineering And Computer Sciences, vol. 4, no. 9, pp. 251-257, 2025.
P. Sharma, “Advancements in OCR: a deep learning algorithm for enhanced text recognition,” International Journal of Inventive Engineering and Sciences, vol. 10, no. 8, pp. 1-7, 2023, doi: 10.35940/ijies.F4263.0810823.
K. Soni V, V. Shukla, R. Tandan S, et al., “Performance Evaluation of Efficient and Accurate Text Detection and Recognition in Natural Scenes Images Using EAST and OCR Fusion,” International Journal of Advanced Computer Science & Applications, vol. 16, no. 1, pp. 445-453, 2025, doi: 10.14569/IJACSA.2025.0160144.
H. Gbada, K. Kalti, and A. Mahjoub M, “Deep learning approaches for information extraction from visually rich documents: datasets, challenges and methods,” International Journal on Document Analysis and Recognition (IJDAR), vol. 28, no. 1, pp. 121-142, 2025, doi: 10.1007/s10032-024-00493-8.
F. Wang, S. Tian, L. Yu, et al., “TEDT: transformer-based encoding-decoding translation network for multimodal sentiment analysis,” Cognitive Computation, vol. 15, no. 1, pp. 289-303, 2023, doi: 10.1007/s12559-022-10073-9.
L. Sun, Z. Lian, B. Liu, et al., “Efficient multimodal transformer with dual-level feature restoration for robust multimodal sentiment analysis,” IEEE Transactions on Affective Computing, vol. 15, no. 1, pp. 309-325, 2023, doi: 10.1109/TAFFC.2023.3274829.
P. Xu, X. Zhu, and A. Clifton D, “Multimodal learning with transformers: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 10, pp. 12113-12132, 2023, doi: 10.1109/TPAMI.2023.3275156.
M. Binte Rashid, S. Rahaman M, and P. Rivas, “Navigating the multimodal landscape: A review on integration of text and image data in machine learning architectures,” Machine Learning and Knowledge Extraction, vol. 6, no. 3, pp. 1545-1563, 2024, doi: 10.3390/make6030074.
F. Shi, R. Gao, W. Huang, et al., “Dynamic mdetr: A dynamic multimodal transformer decoder for visual grounding,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 2, pp. 1181-1198, 2023, doi: 10.1109/TPAMI.2023.3328185.
J. Li, Y. Wang, and D. Zhao, “Layer-wise enhanced transformer with multi-modal fusion for image caption,” Multimedia Systems, vol. 29, no. 3, pp. 1043-1056, 2023, doi: 10.1007/s00530-022-01036-z.
H. Zhang, B. Wu, X. Yuan, et al., “Trustworthy graph neural networks: Aspects, methods, and trends,” Proceedings of the IEEE, vol. 112, no. 2, pp. 97-139, 2024, doi: 10.1109/JPROC.2024.3369017.
B. Abimbola, E. de La Cal Marin, and Q. Tan, “Enhancing legal sentiment analysis: A convolutional neural network-long short-term memory document-level model,” Machine Learning and Knowledge Extraction, vol. 6, no. 2, pp. 877-897, 2024, doi: 10.3390/make6020041.
A. Yusuf A, F. Chong, and M. Xianling, “An analysis of graph convolutional networks and recent datasets for visual question answering,” Artificial Intelligence Review, vol. 55, no. 8, pp. 6277-6300, 2022, doi: 10.1007/s10462-022-10151-2.
X. Liu, C. Miao, G. Fiumara, et al., “Information propagation prediction based on spatial-temporal attention and heterogeneous graph convolutional networks,” IEEE Transactions on Computational Social Systems, vol. 11, no. 1, pp. 945-958, 2023, doi: 10.1109/TCSS.2023.3244573.