Di Buó, Christian
(2026)
Beyond reproducibility in AI for lung cancer screening.
[Laurea magistrale], Università di Bologna, Corso di Studio in
Artificial intelligence [LM-DM270], Documento ad accesso riservato.
Documenti full-text disponibili:
Abstract
Deep learning has reshaped medical image analysis, and lung cancer screening is among its most consequential frontiers, with models like Sybil forcasting lung cancer incidence up to
6 years ahead from a single low-dose CT scan.
Before such models can reach clinical deployment, a question arises: can they be trusted, or are predictions merely an artifact of a particular training run?
Building on Sybil and the National Lung Screening Trial (NLST) cohort, this thesis investigates predictive accuracy and clinical trustworthiness together,
asking whether it not only discriminate future risk, but stays scalable, reliable, and reproducible under stochastic training.
This work spans six axes: a CT-BIDS-compliant NIfTI pipeline resolving distributed I/O bottlenecks; multi-node scaling; a Bland-Altman framework quantifying participant-level repeatability across model ensembles; loss and architectural ablations;
a data-volume sensitivity analysis isolating the model's dependence on positive versus negative training cases; and coregistration testing whether spatial priors can maximize expert annotation.
The findings reveal a dichotomy between what metrics report and what participants experience. The model achieved strong average discrimination, but repeatability told a different story:
disagreement between ensembles grew with predicted risk, so the highest-risk participants are those for whom the model is least reproducible. A second issue emerged from infrastructure:
accuracy resisted architecture and loss changes, yet collapsed under poor network topology and positive-case scarcity, showing bottlenecks lie in engineering and curation, not model design.
A coregistered template with a population-level annotation map reached strong predictive performance.
The thesis concludes clinical translation cannot rest on discrimination alone: accuracy is necessary but insufficient for trust, earned participant by participant, once the average is set aside.
Abstract
Deep learning has reshaped medical image analysis, and lung cancer screening is among its most consequential frontiers, with models like Sybil forcasting lung cancer incidence up to
6 years ahead from a single low-dose CT scan.
Before such models can reach clinical deployment, a question arises: can they be trusted, or are predictions merely an artifact of a particular training run?
Building on Sybil and the National Lung Screening Trial (NLST) cohort, this thesis investigates predictive accuracy and clinical trustworthiness together,
asking whether it not only discriminate future risk, but stays scalable, reliable, and reproducible under stochastic training.
This work spans six axes: a CT-BIDS-compliant NIfTI pipeline resolving distributed I/O bottlenecks; multi-node scaling; a Bland-Altman framework quantifying participant-level repeatability across model ensembles; loss and architectural ablations;
a data-volume sensitivity analysis isolating the model's dependence on positive versus negative training cases; and coregistration testing whether spatial priors can maximize expert annotation.
The findings reveal a dichotomy between what metrics report and what participants experience. The model achieved strong average discrimination, but repeatability told a different story:
disagreement between ensembles grew with predicted risk, so the highest-risk participants are those for whom the model is least reproducible. A second issue emerged from infrastructure:
accuracy resisted architecture and loss changes, yet collapsed under poor network topology and positive-case scarcity, showing bottlenecks lie in engineering and curation, not model design.
A coregistered template with a population-level annotation map reached strong predictive performance.
The thesis concludes clinical translation cannot rest on discrimination alone: accuracy is necessary but insufficient for trust, earned participant by participant, once the average is set aside.
Tipologia del documento
Tesi di laurea
(Laurea magistrale)
Autore della tesi
Di Buó, Christian
Relatore della tesi
Correlatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
Deep Learning, Reproducibility, Low-Dose CT, Lung Cancer Screening, Model Robustness, Repeatability, Sybil
Data di discussione della Tesi
21 Luglio 2026
URI
Altri metadati
Tipologia del documento
Tesi di laurea
(NON SPECIFICATO)
Autore della tesi
Di Buó, Christian
Relatore della tesi
Correlatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
Deep Learning, Reproducibility, Low-Dose CT, Lung Cancer Screening, Model Robustness, Repeatability, Sybil
Data di discussione della Tesi
21 Luglio 2026
URI
Gestione del documento: