DiffStyler: Controllable Dual Diffusion for Text-Driven Image Stylization.

Huang, Nisha; Zhang, Yuxin; Tang, Fan; Ma, Chongyang; Huang, Haibin; Dong, Weiming; Xu, Changsheng

Huang, Nisha; Zhang, Yuxin; Tang, Fan; Ma, Chongyang; Huang, Haibin; Dong, Weiming; Xu, Changsheng.

IEEE Trans Neural Netw Learn Syst ; PP2024 Jan 10.

Article em En | MEDLINE | ID: mdl-38198263

ABSTRACT

ABSTRACT

Despite the impressive results of arbitrary image-guided style transfer methods, text-driven image stylization has recently been proposed for transferring a natural image into a stylized one according to textual descriptions of the target style provided by the user. Unlike the previous image-to-image transfer approaches, text-guided stylization progress provides users with a more precise and intuitive way to express the desired style. However, the huge discrepancy between cross-modal inputs/outputs makes it challenging to conduct text-driven image stylization in a typical feed-forward CNN pipeline. In this article, we present DiffStyler, a dual diffusion processing architecture to control the balance between the content and style of the diffused results. The cross-modal style information can be easily integrated as guidance during the diffusion process step-by-step. Furthermore, we propose a content image-based learnable noise on which the reverse denoising process is based, enabling the stylization results to better preserve the structure information of the content image. We validate the proposed DiffStyler beyond the baseline methods through extensive qualitative and quantitative experiments. The code is available at https//github.com/haha-lisa/Diffstyler.

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Tipo de estudo: Qualitative_research Idioma: En Revista: IEEE Trans Neural Netw Learn Syst Ano de publicação: 2024 Tipo de documento: Article

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google