Marangon, Luca
(2026)
Analyzing CLIP’s Limits: Mitigating Its Problems
and Defining Its Structural Boundaries.
[Laurea], Università di Bologna, Corso di Studio in
Informatica [L-DM270]
Documenti full-text disponibili:
Abstract
CLIP (Contrastive Language-Image Pre-training) is a vision-language model first developed by OpenAI in 2021. It immediately gained great popularity thanks to its zero-shot
capabilities and structural simplicity, becoming the most widespread tool for image-text
association. As the model’s architecture is open source, countless variations have been
created over the years, and many papers have focused on studying the capabilities of this
instrument. In this Thesis, the objective is to analyze the limitations and weaknesses of
this incredible tool, as well as some of the attempts made to improve it.
First, we will describe some of the most commonly encountered limitations of CLIP,
such as its difficulty in recognizing relationships between objects, misinterpretation of
negation in prompts, inability to count with precision, and difficulty in capturing the
global structure (such as recognizing the image’s style).
Then the Thesis will explore some of the possible solutions developed in the years, citing
a few of the existing Open models implementing them.
On a final note, we will discuss the possible structural limitation intrinsic to CLIP, noted
by recent articles, that propose the necessity of considerable structural changes in the
model.
Abstract
CLIP (Contrastive Language-Image Pre-training) is a vision-language model first developed by OpenAI in 2021. It immediately gained great popularity thanks to its zero-shot
capabilities and structural simplicity, becoming the most widespread tool for image-text
association. As the model’s architecture is open source, countless variations have been
created over the years, and many papers have focused on studying the capabilities of this
instrument. In this Thesis, the objective is to analyze the limitations and weaknesses of
this incredible tool, as well as some of the attempts made to improve it.
First, we will describe some of the most commonly encountered limitations of CLIP,
such as its difficulty in recognizing relationships between objects, misinterpretation of
negation in prompts, inability to count with precision, and difficulty in capturing the
global structure (such as recognizing the image’s style).
Then the Thesis will explore some of the possible solutions developed in the years, citing
a few of the existing Open models implementing them.
On a final note, we will discuss the possible structural limitation intrinsic to CLIP, noted
by recent articles, that propose the necessity of considerable structural changes in the
model.
Tipologia del documento
Tesi di laurea
(Laurea)
Autore della tesi
Marangon, Luca
Relatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
CLIP,HARD NEGATIVES,AI
Data di discussione della Tesi
15 Luglio 2026
URI
Altri metadati
Tipologia del documento
Tesi di laurea
(NON SPECIFICATO)
Autore della tesi
Marangon, Luca
Relatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
CLIP,HARD NEGATIVES,AI
Data di discussione della Tesi
15 Luglio 2026
URI
Statistica sui download
Gestione del documento: