Multimodal Deep Learning For Eye Disease Recognition From Fundus And OCT Images With Synthetic OCT Generation
Main Article Content
Abstract
Multimodal learning that integrates Color Fundus Photography (CFP) and Optical Coherence Tomography (OCT) can enhance retinal disease recognition by combining surface appearance with depth-resolved structural cues. However, the limited availability of synchronized 2D-3D CFP-OCT datasets hinders the deployment of multimodal AI in real-world screening settings where OCT devices are often unavailable. This study investigates whether synthetic 3D OCT volumes can serve as a practical auxiliary modality to support multimodal classification in fundus-only scenarios. We synthesize 3D OCT from fundus images using three generative paradigms-a 3D GAN (DCGAN/WGAN-GP style), a 3D VAEGAN, and a fundus-conditioned latent diffusion model (3D-DDPM with classifier-free guidance)-and integrate the generated volumes into an EyeMoST- based uncertainty-aware fusion framework. We evaluate the pipeline on RFMiD and ODIR-5K across multiple backbone configurations. Empirically, diffusion-generated OCT yields the most stable fusion performance. In contrast, synthetic OCT does not consistently outperform strong fundus-only models due to domain gap and imperfect biomarker preservation.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
References
World Health Organization, “World report on vision,” WHO, Geneva, Switzerland, 2019.
M. Gandor et al., “Diagnostics of diabetic retinopathy based on fundus photos using machine learning methods with advanced feature engineering algorithms,” Scientific Reports, vol. 15, p. 34486, 2025.
M. Elkholy and M. A. Marzouk, “Deep learning-based classification of eye diseases using convolutional neural network for OCT images,” Frontiers in Computer Science, vol. 5, p. 1252295, 2024.
Y. Li et al., “Multimodal information fusion for glaucoma and diabetic retinopathy classification,” in Ophthalmic Medical Image Analysis, pp. 53–62, 2022.
E. S¨ukei et al., “Multi-modal representation learning in retinal imaging using self-supervised learning for enhanced clinical predictions,” Scientific Reports, vol. 14, p. 26802, 2024.
M. A. Rodr´ıguez, H. AlMarzouqi and P. Liatsis, “Multi-Label Retinal Disease Classification Using Transformers,” in IEEE Journal of Biomedical and Health Informatics, vol. 27, no. 6, pp. 2739-2750, June 2023.
T. K. Yoo et al., “The possibility of the combination of OCT and fundus images for improving the diagnostic accuracy of deep learning for age-related macular degeneration: A preliminary experiment,” Medical & Biological engineering & Computing, vol. 57, no. 3, pp. 677–687, 2019.
K. Zou et al., “Reliable multimodality eye disease screening via mixture of Student’s t distributions,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2023 , pp. 596–606, 2023.
S. Pachade et al., “Retinal fundus multi-disease image dataset (RFMiD): A dataset for multidisease detection research,” Data, vol. 6, no. 2, p. 14, 2021.
N. Li et al., “A benchmark of ocular disease intelligent recognition: One shot for multi-disease detection (ODIR-5K),” in Proc. Bench, pp. 177– 193, 2021.
K. He, X. Zhang, S. Ren and J. Sun, “Deep Residual Learning for Image Recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, pp. 770-778, 2016.
M. Tan and Q. V. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proceedings of the 36th International Conference on Machine Learning, pp. 6105– 6114, 2019.
A. Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” arXiv preprint arXiv: 2010.11929, 2021.
A. Hatamizadeh et al., “UNETR: Transformers for 3D Medical Image Segmentation,” 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, pp. 1748-1758, 2022.
B. Liu et al., “APTOS-2024 challenge report: Generation of synthetic 3D OCT images from fundus photographs,” arXiv preprint arXiv: 2506.07542, 2025.
Z. Wang et al., “Synthetic artificial intelligence using generative adversarial network for retinal imaging in detection of age-related macular degeneration,” Frontiers in Medicine, vol. 10, p. 1184892, 2023.
Y. Xue et al., “Gen-3Diffusion: Realistic imageto-3D generation via 2D and 3D diffusion synergy,” arXiv preprint arXiv: 2412.06698, 2024.
U. Bhimavarapu, N. Chintalapudi and G. Battineni, “Automatic detection and classification of diabetic retinopathy using the improved pooling function in the convolution neural network,” Diagnostics, vol. 13, no. 15, p. 2606, 2023.
R. K. Rasel et al., “Assessing the efficacy of 2D and 3D CNN algorithms in OCT-based glaucoma detection,” Scientific Reports, vol. 14, p. 11758, 2024.