Opening book details…
Can I read LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models on EtoBox?
LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models by Qin, Zhenyue; Yin, Yu; Campbell, Dylan; Wu, Xuansheng; Zou, Ke; Tham, Yih-Chung; Liu, Ninghao; Zhang, Xiuzhen; Chen, Qingyu is a scholarly article available to read on EtoBox.
What is LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models about?
The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potential to assist in understanding anatomical information, diagnosing eye diseases, and drafting interpretations and follow-up plans, thereby reducing the burden on clinicians and improving access to eye care. However, limited benchmarks are available to assess LVLMs' performance in ophthalmology-specific applications. In this study, we introduce LMOD, a large-scale multimodal ophthalmology benchmark consisting of 21,993 instances across (1) five ophthalmic imaging modalities: optical coherence tomography, color fundus photographs, scanning laser ophthalmoscopy, lens photographs, and surgical scenes; (2) free-text, demographic, and disease biomarker information; and (3) primary ophthalmology-specific applications such as anatomical information understanding, disease diagnosis, and subgroup analysis. In addition, we benchmarked 13 state-of-the-art LVLM representatives from closed-source, open-source, and medical domains. The results demonstrate a significant performance d
- Author
- Qin, Zhenyue; Yin, Yu; Campbell, Dylan; Wu, Xuansheng; Zou, Ke; Tham, Yih-Chung; Liu, Ninghao; Zhang, Xiuzhen; Chen, Qingyu
- Published
- 2024
- Language
- EN