Student Emotion Recognition from Low-Quality Videos Using Multimodal Deep Learning

ANDI MAWADDA TAIBA MAWADDA TAIBA; Rizki Yusliana Bakti; Muhammad Faisal; Muhammad Syafaat S. Kuba; Lukman Anas; Emil Agusalim H. T; Fahrim I. Rahman

doi:10.20895/infotel.v18i1.1523

view PDF

Published May 4, 2026

DOI https://doi.org/10.20895/infotel.v18i1.1523

ANDI MAWADDA TAIBA MAWADDA TAIBA

Universitas Muhammadiyah Makassar, Indonesia

Rizki Yusliana Bakti

Universitas Muhammadiyah Makassar, Indonesia

Muhammad Faisal

Universitas Muhammadiyah Makassar, Indonesia

Muhammad Syafaat S. Kuba

Universitas Muhammadiyah Makassar, Indonesia

Lukman Anas

Universitas Muhammadiyah Makassar, Indonesia

Emil Agusalim H. T

Universitas Muhammadiyah Makassar, Indonesia

Fahrim I. Rahman

Universitas Muhammadiyah Makassar, Indonesia

Abstract

Emotion recognition plays a critical role in intelligent e-learning systems by enabling adaptive feedback and timely pedagogical interventions based on students’ affective states. However, most existing approaches rely heavily on visual facial cues, which are highly vulnerable to real-world conditions such as low-resolution video, partial facial occlusion, poor lighting, and unstable network connections commonly encountered in online learning environments. These limitations significantly degrade the performance of unimodal deep learning models. To address this challenge, this study proposes a multimodal deep learning framework for student emotion recognition that is robust to low-quality and occluded video input. The proposed model integrates visual and audio modalities through a hybrid architecture, combining a lightweight CNN-based visual feature extractor with a BiLSTM-based speech emotion model. An attention-based fusion mechanism is employed to adaptively weight cross-modal features, allowing the system to compensate for degraded or missing visual information using complementary acoustic cues. Experimental evaluations are conducted using publicly available datasets representative of realistic online learning scenarios, including DAiSEE and RAVDESS, with additional augmentation to simulate varying levels of occlusion and video degradation. The results demonstrate that the multimodal approach consistently outperforms unimodal baselines, particularly under high occlusion conditions, while maintaining computational efficiency suitable for near real-time deployment. These findings confirm that multimodal fusion with attention mechanisms provides a more resilient and practical solution for emotion-aware e-learning systems operating under non-ideal input conditions

Downloads

Download data is not yet available.

How to Cite

[1]

A. M. T. TAIBA, “Student Emotion Recognition from Low-Quality Videos Using Multimodal Deep Learning”, INFOTEL, vol. 18, no. 1, pp. 157-179, May 2026.

Issue

Vol 18 No 1 (2026): February

Section

Informatics

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

Authors who publish with this journal agree to the following terms:

Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work

Article Sidebar

Main Article Content

Abstract

Downloads

Article Details