Unimodal Speech Emotion Recognition with Signal-to-Image Transformation and Deep Learning on the KTU-MEDAFE Dataset


Mumcu B., Hatipoğlu Yılmaz B.

4th Cognitive Models and Artificial Intelligence Conference, AICCONF 2026, Prague, Çek Cumhuriyeti, 24 - 25 Nisan 2026, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/aicconf69182.2026.11600663
  • Basıldığı Şehir: Prague
  • Basıldığı Ülke: Çek Cumhuriyeti
  • Anahtar Kelimeler: emotion recognition, KTUMEDAFE, speech signals
  • Karadeniz Teknik Üniversitesi Adresli: Evet

Özet

Emotions are a central aspect of human experience, profoundly influencing communication, decision-making, and social interaction. In this study, we employed the Karadeniz Technical University Multimodal Emotion Dataset using Audio, Facial Images, and EEG (KTU-MEDAFE), focusing on the speech recordings collected while participants read text-based emotional stimuli. The recorded audio signals were converted into spectrogram-like images using a signal-to-image transformation technique, allowing the application of deep neural networks for visual-based analysis. These images were subsequently used to train and evaluate models for speech-based emotion recognition, considering the four emotional categories defined in the dataset: funny, surprising, sad, and neutral. This approach demonstrates the potential of utilizing unimodal speech signals extracted from a multimodal corpus to achieve reliable and robust emotion classification.