|
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
|
| Volume 187 - Issue 135 |
| Published: August 2026 |
| Authors: Sudhi S., Krishnaja M.K. |
10.5120/ijca061ce9fe9af8
|
Sudhi S., Krishnaja M.K. . A Proposed Spectrogram-based Deep Learning Framework for Automatic Swara and Gamaka Recognition in Carnatic Veena Performances. International Journal of Computer Applications. 187, 135 (August 2026), 51-55. DOI=10.5120/ijca061ce9fe9af8
@article{ 10.5120/ijca061ce9fe9af8,
author = { Sudhi S.,Krishnaja M.K. },
title = { A Proposed Spectrogram-based Deep Learning Framework for Automatic Swara and Gamaka Recognition in Carnatic Veena Performances },
journal = { International Journal of Computer Applications },
year = { 2026 },
volume = { 187 },
number = { 135 },
pages = { 51-55 },
doi = { 10.5120/ijca061ce9fe9af8 },
publisher = { Foundation of Computer Science (FCS), NY, USA }
}
%0 Journal Article
%D 2026
%A Sudhi S.
%A Krishnaja M.K.
%T A Proposed Spectrogram-based Deep Learning Framework for Automatic Swara and Gamaka Recognition in Carnatic Veena Performances%T
%J International Journal of Computer Applications
%V 187
%N 135
%P 51-55
%R 10.5120/ijca061ce9fe9af8
%I Foundation of Computer Science (FCS), NY, USA
Music Information Retrieval (MIR) has witnessed significant advances [1], [8]; with the adoption of artificial intelligence and deep learning techniques; however, the computational analysis of Indian classical music remains a challenging research problem because of its intricate melodic structure, microtonal variations, and expressive ornamentations. Among Carnatic musical instruments, the Saraswati Veena presents additional challenges due to its sustained resonance, continuous pitch transitions, and complex playing techniques that produce highly dynamic acoustic patterns. Existing approaches have primarily focused on pitch estimation, raga recognition, or instrument classification, while the automatic recognition of swaras and gamakas in Veena performances has received comparatively limited attention. This study proposes a spectrogram-based deep learning framework for the automatic recognition of swaras and gamakas from Carnatic Veena recordings. The proposed methodology incorporates audio acquisition, preprocessing, silence removal, noise reduction, segmentation, normalization, and data augmentation to improve the quality and consistency of audio samples. The processed signals are transformed into Short-Time Fourier Transform (STFT) and Mel spectrogram representations, which effectively preserve temporal and spectral information for feature extraction. These spectrograms are used as inputs to a Convolutional Neural Network (CNN) that automatically learns discriminative acoustic features for multi-class classification. A representative dataset specification consisting of seven swara categories and five gamaka categories is included to demonstrate the proposed computational workflow and provide a reproducible framework for future investigations. The proposed framework has potential applications in automatic music transcription, intelligent music tutoring, performer assistance, digital preservation of Carnatic musical heritage, and Music Information Retrieval. Future research will focus on validating the framework using expert-annotated datasets and enhancing recognition performance through transformer-based architectures, attention mechanisms, transfer learning, and multimodal learning techniques.