A Multimodal Framework for Misinformation Detection in Bengali Using Large Language Models

Sajida Kabir, Mahadi Hasan, Abdulla Masud, John Gomes, Nafees Mansoor

Conference Paper. 2025 IEEE 4th International Conference on Robotics, Automation, Artificial-Intelligence and Internet-of-Things, RAAICON 2025, pp. 444–449 (2025).

Abstract

Misinformation has emerged as a pressing societal challenge in the digital era, with its impact particularly severe in low-resource linguistic contexts such as Bengali, where detection tools and datasets remain limited. This paper addresses this gap through two key contributions. First, we present BD-FakeDetect, a novel, manually curated multimodal dataset comprising 5,929 Bengali news articles collected from verified fact-checking platforms and annotated across eight finegrained deception categories. Second, we propose a robust detection framework built upon a fine-tuned GPT4o model, a state-of-the-art large language model with native multimodal capabilities. To enhance factual grounding and mitigate hallucination, the framework integrates a Retrieval-Augmented Generation (RAG) pipeline that incorporates external evidence in real time. Empirical evaluation underscores the framework's effectiveness: on the complex multimodal detection task, the final fine-tuned GPT-4o model achieved a validation accuracy of 81.8%. This was benchmarked against a text-only evaluation, which established a strong baseline with the DeepSeek 7B model achieving 90% token accuracy. Collectively, this work establishes a comprehensive benchmark for multimodal factchecking in the Bengali digital landscape and advances culturally and linguistically grounded research on misinformation. © 2025 IEEE.

Keywords

Bengali NLP, Computer Vision, Deepfake, Fact-Checking, Forgery detection, Large Language Models, Low-Resource Languages, Misinformation Detection, Multimodal Learning, Splice detection

DOI: 10.1109/raaicon69033.2025.11502571