Omelchenko, Maksim
(2026)
Improving prompt adherence in identity-preserving text-to-image generative models.
[Laurea magistrale], Università di Bologna, Corso di Studio in
Artificial intelligence [LM-DM270], Documento full-text non disponibile
Il full-text non è disponibile per scelta dell'autore.
(
Contatta l'autore)
Abstract
Text-to-image diffusion models can generate highly realistic images from natural language prompts, but controlling semantic attributes while preserving subject identity remains difficult. Identity-preserving adapters such as PuLID maintain strong identity consistency but often reduce responsiveness to prompt-driven changes in attributes such as age, gender, or facial expression.
This thesis proposes a method for improving semantic controllability by manipulating identity embeddings directly in PuLID space. Semantic direction vectors are constructed by computing centroid differences between embeddings of images belonging to different attribute classes. These directions are applied to the identity embedding before generation, enabling controlled semantic transformations while retaining core identity features. To improve robustness, a variance-based filtering strategy is introduced to suppress unstable embedding components.
The method is evaluated on multiple facial attribute edits, including age, gender, and expression changes, using CLIP-direction scores to measure prompt alignment and ArcFace similarity to assess identity preservation. Results show that embedding manipulation substantially improves semantic alignment compared to prompt-only editing, while variance filtering reduces identity drift. Overall, the findings demonstrate that structured semantic directions exist in PuLID embedding space and can be exploited to enable controllable identity-conditioned image generation.
Abstract
Text-to-image diffusion models can generate highly realistic images from natural language prompts, but controlling semantic attributes while preserving subject identity remains difficult. Identity-preserving adapters such as PuLID maintain strong identity consistency but often reduce responsiveness to prompt-driven changes in attributes such as age, gender, or facial expression.
This thesis proposes a method for improving semantic controllability by manipulating identity embeddings directly in PuLID space. Semantic direction vectors are constructed by computing centroid differences between embeddings of images belonging to different attribute classes. These directions are applied to the identity embedding before generation, enabling controlled semantic transformations while retaining core identity features. To improve robustness, a variance-based filtering strategy is introduced to suppress unstable embedding components.
The method is evaluated on multiple facial attribute edits, including age, gender, and expression changes, using CLIP-direction scores to measure prompt alignment and ArcFace similarity to assess identity preservation. Results show that embedding manipulation substantially improves semantic alignment compared to prompt-only editing, while variance filtering reduces identity drift. Overall, the findings demonstrate that structured semantic directions exist in PuLID embedding space and can be exploited to enable controllable identity-conditioned image generation.
Tipologia del documento
Tesi di laurea
(Laurea magistrale)
Autore della tesi
Omelchenko, Maksim
Relatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
text-to-image generation, diffusion models, identity preservation, controllable image generation, embedding manipulation, facial attribute editing, generative AI
Data di discussione della Tesi
26 Marzo 2026
URI
Altri metadati
Tipologia del documento
Tesi di laurea
(NON SPECIFICATO)
Autore della tesi
Omelchenko, Maksim
Relatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
text-to-image generation, diffusion models, identity preservation, controllable image generation, embedding manipulation, facial attribute editing, generative AI
Data di discussione della Tesi
26 Marzo 2026
URI
Gestione del documento: