Research Article

A Proposed Spectrogram-based Deep Learning Framework for Automatic Swara and Gamaka Recognition in Carnatic Veena Performances

by  Sudhi S., Krishnaja M.K.
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 135
Published: August 2026
Authors: Sudhi S., Krishnaja M.K.
10.5120/ijca061ce9fe9af8
PDF

Sudhi S., Krishnaja M.K. . A Proposed Spectrogram-based Deep Learning Framework for Automatic Swara and Gamaka Recognition in Carnatic Veena Performances. International Journal of Computer Applications. 187, 135 (August 2026), 51-55. DOI=10.5120/ijca061ce9fe9af8

                        @article{ 10.5120/ijca061ce9fe9af8,
                        author  = { Sudhi S.,Krishnaja M.K. },
                        title   = { A Proposed Spectrogram-based Deep Learning Framework for Automatic Swara and Gamaka Recognition in Carnatic Veena Performances },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 135 },
                        pages   = { 51-55 },
                        doi     = { 10.5120/ijca061ce9fe9af8 },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Sudhi S.
                        %A Krishnaja M.K.
                        %T A Proposed Spectrogram-based Deep Learning Framework for Automatic Swara and Gamaka Recognition in Carnatic Veena Performances%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 135
                        %P 51-55
                        %R 10.5120/ijca061ce9fe9af8
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Music Information Retrieval (MIR) has witnessed significant advances [1], [8]; with the adoption of artificial intelligence and deep learning techniques; however, the computational analysis of Indian classical music remains a challenging research problem because of its intricate melodic structure, microtonal variations, and expressive ornamentations. Among Carnatic musical instruments, the Saraswati Veena presents additional challenges due to its sustained resonance, continuous pitch transitions, and complex playing techniques that produce highly dynamic acoustic patterns. Existing approaches have primarily focused on pitch estimation, raga recognition, or instrument classification, while the automatic recognition of swaras and gamakas in Veena performances has received comparatively limited attention. This study proposes a spectrogram-based deep learning framework for the automatic recognition of swaras and gamakas from Carnatic Veena recordings. The proposed methodology incorporates audio acquisition, preprocessing, silence removal, noise reduction, segmentation, normalization, and data augmentation to improve the quality and consistency of audio samples. The processed signals are transformed into Short-Time Fourier Transform (STFT) and Mel spectrogram representations, which effectively preserve temporal and spectral information for feature extraction. These spectrograms are used as inputs to a Convolutional Neural Network (CNN) that automatically learns discriminative acoustic features for multi-class classification. A representative dataset specification consisting of seven swara categories and five gamaka categories is included to demonstrate the proposed computational workflow and provide a reproducible framework for future investigations. The proposed framework has potential applications in automatic music transcription, intelligent music tutoring, performer assistance, digital preservation of Carnatic musical heritage, and Music Information Retrieval. Future research will focus on validating the framework using expert-annotated datasets and enhancing recognition performance through transformer-based architectures, attention mechanisms, transfer learning, and multimodal learning techniques.

References
  • H. Purwins, B. Li, T. Virtanen, J. Schluter, S.-Y. Chang, and T. Sainath, "Deep learning for audio signal processing," IEEE Journal of Selected Topics in Signal Processing, vol. 13, no. 2, pp. 206–219, 2019. DOI: https://doi.org/10.1109/JSTSP.2019.2908700
  • S. R. M. Prasanna and P. Dhanalakshmi, "Automatic recognition of Indian classical music using machine learning techniques: A survey," Journal of Intelligent Information Systems, vol. 58, no. 2, pp. 415–442, 2022.
  • Y. Gong, Y.-A. Chung, and J. Glass, "AST: Audio Spectrogram Transformer," Proceedings of Interspeech, 2021.
  • Y. Gong et al., "SSAST: Self-Supervised Audio Spectrogram Transformer," Proceedings of AAAI Conference on Artificial Intelligence, 2022. DOI: https://doi.org/ 10.1609/aaai.v36i12.21452
  • A. Baevski, H. Zhou, A. Mohamed, and M. Auli, "wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations," Advances in Neural Information Processing Systems, 2020.
  • A. Gulati et al., "Conformer: Convolution-Augmented Transformer for Speech Recognition," Proceedings of Interspeech, 2020. DOI: https://doi.org/ 10.21437/Interspeech.2020-3015
  • K. Choi, G. Fazekas, and M. Sandler, "The Effects of Noise and Data Augmentation on Deep Neural Networks for Music Classification," Applied Sciences, vol. 10, no. 3, 2020. DOI: https://doi.org/ 10.3390/app10030908
  • J. Pons and X. Serra, "Deep Learning for Music Information Retrieval: Challenges and Directions," IEEE Signal Processing Magazine, vol. 38, no. 6, pp. 12–23, 2021. DOI: https://doi.org/ 10.1109/MSP.2021.3115817
  • M. Won, S. Chun, O. Nieto, and X. Serra, "Data-driven harmonic analysis of music using deep neural networks," IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 2889–2900, 2021. DOI: https://doi.org/ 10.1109/TASLP.2021.3110131
  • E. Fonseca et al., "Learning Sound Event Classification from Weakly Labeled Data," IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3218–3231, 2021. DOI: https://doi.org/ 10.1109/TASLP.2021.3119290
  • H. Chen and J. Yang, "Spectrogram-based audio classification using convolutional neural networks: A review," Electronics, vol. 10, no. 18, 2021. DOI: https://doi.org/ 10.3390/electronics10182231
  • S. Hershey et al., "CNN architectures for large-scale audio classification," IEEE ICASSP, 2017. (Foundational reference) DOI: https://doi.org/ 10.1109/ICASSP.2017.7952132
  • F. Chollet, Deep Learning with Python, 2nd ed., Manning Publications, 2021.
  • A. Dosovitskiy et al., "An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale," ICLR, 2021.
  • Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, "HTS-AT: A Hierarchical Token-Semantic Audio Transformer," ICASSP, 2022. DOI: https://doi.org/ 10.1109/ICASSP43922.2022.9746126
  • J. Lee, J. Park, K. Kim, and S. Choi, "Transfer Learning for Environmental Sound Classification Using Spectrogram-Based CNNs," Sensors, vol. 22, 2022. DOI: https://doi.org/ 10.3390/s22093280
  • X. Li, Y. Wang, and H. Wang, "Audio classification using Vision Transformers and Mel Spectrograms," Applied Sciences, vol. 13, 2023.
  • S. K. Sharma and R. Gupta, "Deep learning approaches for music information retrieval: A comprehensive review," Multimedia Tools and Applications, vol. 82, 2023.
  • A. Kumar and P. Rao, "Automatic recognition of Carnatic music patterns using deep neural networks," Journal of New Music Research, vol. 52, 2023.
  • V. Narayanan, S. Gulati, and X. Serra, "Computational approaches to Indian classical music analysis," ACM Computing Surveys, vol. 56, no. 2, 2024.
  • S. Gulati and X. Serra, "Deep Learning Methods for Computational Analysis of Carnatic Music," Transactions of the International Society for Music Information Retrieval, 2024.
  • J. Brown, L. Smith, and R. Jones, "Attention-based neural architectures for music transcription," IEEE Access, vol. 12, 2024.
  • K. Patel and M. Singh, "Spectrogram-based music classification using hybrid CNN–Transformer networks," Pattern Recognition Letters, vol. 183, 2025.
  • H. Zhao, Y. Liu, and X. Chen, "Self-supervised learning for audio event recognition: A review," Information Fusion, vol. 106, 2025.
  • S. Verma and A. Iyer, "Deep learning techniques for automatic music transcription and classification: Recent advances and future directions," Expert Systems with Applications, vol. 260, 2025.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Carnatic Music Veena Swara Recognition Gamaka Recognition Convolutional Neural Network Mel Spectrogram Short-Time Fourier Transform Audio Signal Processing

Powered by PhDFocusTM