Thangka Image Captioning Based on Semantic Concept Prompt and Multimodal Feature Optimization.

Hu, Wenjin; Qiao, Lang; Kang, Wendong; Shi, Xinyue

Hu, Wenjin; Qiao, Lang; Kang, Wendong; Shi, Xinyue.

Afiliación

Hu W; School of Mathematics and Computer Science, Northwest Minzu Univsersity, Lanzhou 730030, China.
Qiao L; Key Laboratory of China's Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, China.
Kang W; Key Laboratory of China's Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, China.
Shi X; Key Laboratory of China's Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, China.

J Imaging ; 9(8)2023 Aug 16.

Article en En | MEDLINE | ID: mdl-37623694

ABSTRACT

ABSTRACT

Thangka images exhibit a high level of diversity and richness, and the existing deep learning-based image captioning methods generate poor accuracy and richness of Chinese captions for Thangka images. To address this issue, this paper proposes a Semantic Concept Prompt and Multimodal Feature Optimization network (SCAMF-Net). The Semantic Concept Prompt (SCP) module is introduced in the text encoding stage to obtain more semantic information about the Thangka by introducing contextual prompts, thus enhancing the richness of the description content. The Multimodal Feature Optimization (MFO) module is proposed to optimize the correlation between Thangka images and text. This module enhances the correlation between the image features and text features of the Thangka through the Captioner and Filter to more accurately describe the visual concept features of the Thangka. The experimental results demonstrate that our proposed method outperforms baseline models on the Thangka dataset in terms of BLEU-4, METEOR, ROUGE, CIDEr, and SPICE by 8.7%, 7.9%, 8.2%, 76.6%, and 5.7%, respectively. Furthermore, this method also exhibits superior performance compared to the state-of-the-art methods on the public MSCOCO dataset.

Palabras clave

Thangka; deep learning; image captioning; knowledge distillation; visual concepts

Texto completo

Añadir a Mi BVS

Imprimir

XML

PubMed Links

Buscar en Google

Texto completo: 1 Colección: 01-internacional Base de datos: MEDLINE Idioma: En Revista: J Imaging Año: 2023 Tipo del documento: Article País de afiliación: China