BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalities

Sahal Shaji Mullappilly; Mohammed Irfan Kurpath; Sara Pieri; Saeed Yahya Alseiari; Shanavas Cholakkal; Khaled M Aldahmani; Fahad Shahbaz Khan; Rao Muhammad Anwer; Salman Khan; Timothy Baldwin; Hisham Cholakkal

2025 EMNLP EMNLP 2025

BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalities

Abstract

AbstractWe introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions. It enables multi-turn conversation in Arabic and English and supports diverse medical imaging modalities, including radiology, CT, and histology. To train BiMediX2, we curate BiMed-V, an extensive Arabic-English bilingual healthcare dataset consisting of 1.6M samples of diverse medical interactions. This dataset supports a range of medical Large Language Model (LLM) and Large Multimodal Model (LMM) tasks, including multi-turn medical conversations, report generation, and visual question answering (VQA). We also introduce BiMed-MBench, the first Arabic-English medical LMM evaluation benchmark, verified by medical experts. BiMediX2 demonstrates excellent performance across multiple medical LLM and LMM benchmarks, achieving state-of-the-art results compared to other open-sourced models. On BiMed-MBench, BiMediX2 outperforms existing methods by over 9% in English and more than 20% in Arabic evaluations. Additionally, it surpasses GPT-4 by approximately 9% in UPHILL factual accuracy evaluations and excels in various medical VQA, report generation, and report summarization tasks. Our trained models, instruction set, and source code are available at - https://github.com/mbzuai-oryx/BiMediX2

🌉 Interdisciplinary Bridge — Artificial Intelligence and Deep Learning and Healthcare & Medicine and Natural Language Processing

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Sahal Shaji Mullappilly , Mohammed Irfan Kurpath , Sara Pieri , Saeed Yahya Alseiari , Shanavas Cholakkal , Khaled M Aldahmani , Fahad Shahbaz Khan , Rao Muhammad Anwer , Salman Khan , Timothy Baldwin , Hisham Cholakkal

Topics

Artificial Intelligence > Core AI > Multimodal Learning Natural Language Processing > Resources & Methods > Large Language Models Healthcare & Medicine > Clinical > Medical Imaging Deep Learning > Models > Large Language Models Natural Language Processing > Applications > Visual Question Answering

Keywords

visual question answering medical imaging medical report generation large multimodal model bilingual model large language model

Download PDF

Related papers

Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense Framework 2025

VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing 2025

Model-based Large Language Model Customization as Service 2025

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration 2025

SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design 2025