Enhanced Hybrid Vision Transformer with Multi-Scale Feature Integration and Patch Dropping for Facial Expression Recognition.

Li, Nianfeng; Huang, Yongyuan; Wang, Zhenyan; Fan, Ziyao; Li, Xinyuan; Xiao, Zhiguo

Li, Nianfeng; Huang, Yongyuan; Wang, Zhenyan; Fan, Ziyao; Li, Xinyuan; Xiao, Zhiguo.

Afiliação

Li N; College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China.
Huang Y; College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China.
Wang Z; College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China.
Fan Z; College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China.
Li X; College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China.
Xiao Z; College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China.

Sensors (Basel) ; 24(13)2024 Jun 26.

Article em En | MEDLINE | ID: mdl-39000930

ABSTRACT

ABSTRACT

Convolutional neural networks (CNNs) have made significant progress in the field of facial expression recognition (FER). However, due to challenges such as occlusion, lighting variations, and changes in head pose, facial expression recognition in real-world environments remains highly challenging. At the same time, methods solely based on CNN heavily rely on local spatial features, lack global information, and struggle to balance the relationship between computational complexity and recognition accuracy. Consequently, the CNN-based models still fall short in their ability to address FER adequately. To address these issues, we propose a lightweight facial expression recognition method based on a hybrid vision transformer. This method captures multi-scale facial features through an improved attention module, achieving richer feature integration, enhancing the network's perception of key facial expression regions, and improving feature extraction capabilities. Additionally, to further enhance the model's performance, we have designed the patch dropping (PD) module. This module aims to emulate the attention allocation mechanism of the human visual system for local features, guiding the network to focus on the most discriminative features, reducing the influence of irrelevant features, and intuitively lowering computational costs. Extensive experiments demonstrate that our approach significantly outperforms other methods, achieving an accuracy of 86.51% on RAF-DB and nearly 70% on FER2013, with a model size of only 3.64 MB. These results demonstrate that our method provides a new perspective for the field of facial expression recognition.

Assuntos

Expressão Facial; Redes Neurais de Computação; Humanos; Reconhecimento Facial Automatizado/métodos; Algoritmos; Processamento de Imagem Assistida por Computador/métodos; Face; Reconhecimento Automatizado de Padrão/métodos

Palavras-chave

attention module; facial expression recognition; lightweight network; transformer

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Base de dados: MEDLINE Assunto principal: Redes Neurais de Computação / Expressão Facial Idioma: En Ano de publicação: 2024 Tipo de documento: Article

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Base de dados: MEDLINE Assunto principal: Redes Neurais de Computação / Expressão Facial Idioma: En Ano de publicação: 2024 Tipo de documento: Article