Marchi, Federico
(2026)
Unsupervised Optimization of Accuracy and Interpretability in Multimodal Models for Medical Disease Diagnosis.
[Laurea magistrale], Università di Bologna, Corso di Studio in
Artificial intelligence [LM-DM270], Documento full-text non disponibile
Il full-text non è disponibile per scelta dell'autore.
(
Contatta l'autore)
Abstract
In recent years, contrastive models like CLIP have dominated cross-modal retrieval by aligning images and text in a shared embedding space for zero-shot flexibility. Yet most represent images as single global embeddings, lacking mechanisms to identify regions driving text similarity, a fundamental barrier to interpretable retrieval.
In the medical field, this limitation is critical. Models like GLoRIA, BioViL and BiomedCLIP improve representations via domain-specific pre-training but discard spatial structure. In chest radiography, when pathologies occupy only a small fraction of the image, indiscriminate processing of all patches dilutes discriminative signals and obscures which regions drive predictions.
This thesis systematically explores families of text-conditioned patch selection strategies built on BiomedCLIP, progressively moving from post-ViT local alignment through external patch masking to fully integrated intra-ViT selection, where patches are scored via FILIP (Fine-grained Interactive Language-Image Pre-training) at intermediate layers and either hard-dropped or soft-gated. Each approach is evaluated across multiple configurations on MIMIC-CXR and MS-CXR, measuring retrieval quality, interpretability, and embedding structure.
Results show that text-conditioning is the critical design choice: FILIP-based intra-ViT selection consistently improves embedding structure over the global contrastive baseline without sacrificing retrieval quality. Among intra-ViT methods, soft sigmoid gating achieves the strongest spatial grounding, while adaptive thresholding produces dramatically more faithful patch attributions than any fixed-budget approach. These findings establish text-conditioned intra-ViT patch selection as a promising direction for interpretable vision-language retrieval, revealing a fundamental trade-off between clustering quality, retrieval, and faithfulness that provides concrete design guidelines for future architectures beyond the radiology domain.
Abstract
In recent years, contrastive models like CLIP have dominated cross-modal retrieval by aligning images and text in a shared embedding space for zero-shot flexibility. Yet most represent images as single global embeddings, lacking mechanisms to identify regions driving text similarity, a fundamental barrier to interpretable retrieval.
In the medical field, this limitation is critical. Models like GLoRIA, BioViL and BiomedCLIP improve representations via domain-specific pre-training but discard spatial structure. In chest radiography, when pathologies occupy only a small fraction of the image, indiscriminate processing of all patches dilutes discriminative signals and obscures which regions drive predictions.
This thesis systematically explores families of text-conditioned patch selection strategies built on BiomedCLIP, progressively moving from post-ViT local alignment through external patch masking to fully integrated intra-ViT selection, where patches are scored via FILIP (Fine-grained Interactive Language-Image Pre-training) at intermediate layers and either hard-dropped or soft-gated. Each approach is evaluated across multiple configurations on MIMIC-CXR and MS-CXR, measuring retrieval quality, interpretability, and embedding structure.
Results show that text-conditioning is the critical design choice: FILIP-based intra-ViT selection consistently improves embedding structure over the global contrastive baseline without sacrificing retrieval quality. Among intra-ViT methods, soft sigmoid gating achieves the strongest spatial grounding, while adaptive thresholding produces dramatically more faithful patch attributions than any fixed-budget approach. These findings establish text-conditioned intra-ViT patch selection as a promising direction for interpretable vision-language retrieval, revealing a fundamental trade-off between clustering quality, retrieval, and faithfulness that provides concrete design guidelines for future architectures beyond the radiology domain.
Tipologia del documento
Tesi di laurea
(Laurea magistrale)
Autore della tesi
Marchi, Federico
Relatore della tesi
Correlatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
Vision-Language Alignment, Medical Image Retrieval, Medical AI, Natural Language Processing, Large Language Models
Data di discussione della Tesi
26 Marzo 2026
URI
Altri metadati
Tipologia del documento
Tesi di laurea
(NON SPECIFICATO)
Autore della tesi
Marchi, Federico
Relatore della tesi
Correlatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
Vision-Language Alignment, Medical Image Retrieval, Medical AI, Natural Language Processing, Large Language Models
Data di discussione della Tesi
26 Marzo 2026
URI
Gestione del documento: