Skip to main navigation Skip to search Skip to main content

A multimodal deep learning approach for legal English learning in intelligent educational systems

  • Yanlin Chen
  • , Chenjia Huang
  • , Shumiao Gao
  • , Yifan Lyu
  • , Xinyuan Chen
  • , Shen Liu
  • , Dat Bao
  • , Chunli Lv

Research output: Contribution to journalArticleResearchpeer-review

Abstract

With the development of artificial intelligence and intelligent sensor technologies, traditional legal English teaching approaches have faced numerous challenges in handling multimodal inputs and complex reasoning tasks. In response to these challenges, a cross-modal legal English question-answering system based on visual and acoustic sensor inputs was proposed, integrating image, text, and speech information and adopting a unified vision–language–speech encoding mechanism coupled with dynamic attention modeling to effectively enhance learners’ understanding and expressive abilities in legal contexts. The system exhibited superior performance across multiple experimental evaluations. In the assessment of question-answering accuracy, the proposed method achieved the best results across BLEU, ROUGE, Precision, Recall, and Accuracy, with an Accuracy of 0.87, Precision of 0.88, and Recall of 0.85, clearly outperforming the traditional ASR+SVM classifier, image-retrieval-based QA model, and unimodal BERT QA system. In the analysis of multimodal matching performance, the proposed method achieved optimal results in Matching Accuracy, Recall@1, Recall@5, and MRR, with a Matching Accuracy of 0.85, surpassing mainstream cross-modal models such as VisualBERT, LXMERT, and CLIP. The user study further verified the system’s practical effectiveness in real teaching environments, with learners’ understanding improvement reaching 0.78, expression improvement reaching 0.75, and satisfaction score reaching 0.88, significantly outperforming traditional teaching methods and unimodal systems. The experimental results fully demonstrate that the proposed cross-modal legal English question-answering system not only exhibits significant advantages in multimodal feature alignment and deep reasoning modeling but also shows substantial potential in enhancing learners’ comprehensive capabilities and learning experiences.

Original languageEnglish
Article number3397
Number of pages30
JournalSensors
Volume25
Issue number11
DOIs
Publication statusPublished - 2025

Keywords

  • visual and acoustic sensor integration
  • multimodal semantic fusion
  • human-centered intelligent education
  • vision–language–speech unified encoding

Cite this