Research Article

Multi-Modal Deep Learning Framework for Prediction and Early Identification of Neoplasms: A Hybrid CNN-Transformer Architecture with Explainability

by  Jyotsna Patel, Ankur Pandey, Pushpendra Singh Tomar
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 138
Published: August 2026
Authors: Jyotsna Patel, Ankur Pandey, Pushpendra Singh Tomar
10.5120/ijca69c360f8e2a2
PDF

Jyotsna Patel, Ankur Pandey, Pushpendra Singh Tomar . Multi-Modal Deep Learning Framework for Prediction and Early Identification of Neoplasms: A Hybrid CNN-Transformer Architecture with Explainability. International Journal of Computer Applications. 187, 138 (August 2026), 9-15. DOI=10.5120/ijca69c360f8e2a2

                        @article{ 10.5120/ijca69c360f8e2a2,
                        author  = { Jyotsna Patel,Ankur Pandey,Pushpendra Singh Tomar },
                        title   = { Multi-Modal Deep Learning Framework for Prediction and Early Identification of Neoplasms: A Hybrid CNN-Transformer Architecture with Explainability },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 138 },
                        pages   = { 9-15 },
                        doi     = { 10.5120/ijca69c360f8e2a2 },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Jyotsna Patel
                        %A Ankur Pandey
                        %A Pushpendra Singh Tomar
                        %T Multi-Modal Deep Learning Framework for Prediction and Early Identification of Neoplasms: A Hybrid CNN-Transformer Architecture with Explainability%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 138
                        %P 9-15
                        %R 10.5120/ijca69c360f8e2a2
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Early and accurate detection of neoplasms is a critical challenge in clinical oncology, where diagnostic delays substantially reduce patient survival rates. This paper proposes a novel Multi-Modal Hybrid CNN-Transformer (MMHCT) framework for simultaneous early prediction and identification of neoplasms across five cancer types: breast, lung, brain, prostate, and colorectal. The architecture fuses spatial feature extraction via a modified ResNet-50 backbone with long-range dependency modeling through a Vision Transformer (ViT-B/16) encoder. A Cross-Modal Attention Fusion (CMAF) mechanism integrates heterogeneous inputs, including CT, MRI, whole-slide histopathology images, and structured Electronic Health Records (EHR). A custom focal Dice loss function and federated learning protocol address class imbalance and data privacy constraints, respectively. On the TCGA-LUAD, CBIS-DDSM, BraTS-2023, and Patch Camelyon benchmarks, MMHCT achieves mean AUC = 0.974, sensitivity = 94.3%, specificity = 96.1%, and F1-score = 0.943 at stage-I detection, outperforming all nine baseline methods. GRAD-CAM++ saliency maps and SHAP feature attributions are embedded to ensure clinical interpretability. The framework complies with GDPR and HIPAA regulations through differential privacy mechanisms with privacy budget epsilon = 0.3.

References
  • H. Sung, J. Ferlay, R. L. Siegel et al., “Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality for 36 cancers in 185 countries,” CA: A Cancer Journal for Clinicians, vol. 74, no. 3, pp. 229-263, 2024.
  • American Cancer Society, Cancer Facts & Figures 2024. Atlanta: ACS, 2024.
  • J. G. Elmore, G. M. Longton, P. A. Carney et al., “Variability in pathologists’ interpretations of individual breast biopsy slides,” Annals of Internal Medicine, vol. 164, no. 10, pp. 649-655, 2015.
  • G. Litjens, T. Kooi, B. E. Bejnordi et al., “A survey on deep learning in medical image analysis,” Medical Image Analysis, vol. 42, pp. 60-88, 2017.
  • P. Rajpurkar, J. Irvin, K. Ball et al., “CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning,” arXiv:1711.05225, 2017.
  • A. Esteva, B. Kuprel, R. A. Novoa et al., “Dermatologist-level classification of skin cancer with deep neural networks,” Nature, vol. 542, no. 7639, pp. 115-118, 2017.
  • A. Dosovitskiy, L. Beyer, A. Kolesnikov et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. ICLR, 2021.
  • J. Chen, Y. Lu, Q. Yu et al., “TransUNet: Transformers make strong encoders for medical image segmentation,” arXiv:2102.04306, 2021.
  • Z. Liu, Y. Lin, Y. Cao et al., “Swin Transformer: Hierarchical vision transformer using shifted windows,” in Proc. ICCV, pp. 10012-10022, 2021.
  • P. Mobadersany, S. Yousefi, M. Amgad et al., “Predicting cancer outcomes from histology and genomics using convolutional networks,” Proc. Natl. Acad. Sci., vol. 115, no. 13, pp. E2970-E2979, 2018.
  • R. J. Chen, M. Y. Lu, J. Wang et al., “Multimodal co-attention transformer for survival prediction in gigapixel whole slide images,” in Proc. ICCV, pp. 4015-4025, 2021.
  • R. R. Selvaraju, M. Cogswell, A. Das et al., “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” in Proc. ICCV, pp. 618-626, 2017.
  • S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proc. NeurIPS, vol. 30, 2017.
  • S. Singh, A. Kumar, R. Sharma et al., “Explainability in deep learning for oncology: A systematic review (2020-2024),” npj Digital Medicine, vol. 7, p. 88, 2024.
  • I. Mironov, “Rényi differential privacy,” in Proc. IEEE CSF, pp. 263-275, 2017.
  • K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. CVPR, pp. 770-778, 2016.
  • G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proc. CVPR, pp. 4700-4708, 2017.
  • M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. ICML, pp. 6105-6114, 2019.
  • M. Y. Lu, B. Chen, D. F. K. Williamson et al., “A visual-language foundation model for computational pathology,” Nature Medicine, vol. 30, no. 3, pp. 863-874, 2024.
  • R. J. Chen, C. Ding, M. Y. Lu et al., “Towards a general-purpose foundation model for computational pathology,” Nature Medicine, vol. 30, no. 3, pp. 850-862, 2024.
  • C. D. Lehman, R. D. Arao, B. L. Sprague et al., “National performance benchmarks for modern screening digital mammography,” Radiology, vol. 283, no. 1, pp. 49-58, 2017.
  • E. J. Feuer, B. S. Levy, H. L. Haber et al., “CISNET: Using modeling to understand cancer control,” J. Natl. Cancer Inst. Monogr., no. 56, pp. 2-6, 2020.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Neoplasm detection deep learning Vision Transformer multi-modal fusion federated learning explainable AI oncology medical imaging

Powered by PhDFocusTM