Rezaei, Mohammadhossein
(2026)
Statistical Characterization of Deep Neural Network Intermediate Activations for Adversarial Attacks Detection.
[Laurea magistrale], Università di Bologna, Corso di Studio in
Ingegneria elettronica [LM-DM270]
Documenti full-text disponibili:
Abstract
Deep neural networks have achieved high performance in image classification tasks, yet they remain vulnerable to adversarial examples, where small and often imperceptible input perturbations can lead to incorrect predictions. This thesis investigates whether intermediate activations of convolutional neural networks can be statistically characterized and used for adversarial attack detection.
The proposed approach first extracts layer activations and flattens them into activation vectors. For each selected layer, a layer-wise SVD projection basis is obtained from the corresponding layer transformation. For convolutional layers, this transformation is represented through an equivalent operator form, while linear layers are handled through their weight transformation. Sample activation vectors are then projected onto the corresponding SVD basis to obtain corevectors. Based on these projected representations, a tail-energy statistic is introduced to measure the energy of a sample in the lower singular-value components of the SVD basis. The central hypothesis is that adversarial examples may follow abnormal internal trajectories and therefore produce distinctive energy patterns in selected intermediate layers.
The method is evaluated using standard and adversarially trained Wide Residual Networks under an adversarial robustness framework involving AutoAttack components. The results show that intermediate representations contain useful statistical information for detecting certain adversarial attacks, particularly APGD-based attacks. However, detection remains more difficult for FAB and Square Attack, whose energy distributions overlap more strongly with clean samples. Overall, the study demonstrates that internal activation geometry provides a meaningful direction for adversarial detection, while highlighting the need for broader attack-aware and layer-wise analysis in future robustness research.
Abstract
Deep neural networks have achieved high performance in image classification tasks, yet they remain vulnerable to adversarial examples, where small and often imperceptible input perturbations can lead to incorrect predictions. This thesis investigates whether intermediate activations of convolutional neural networks can be statistically characterized and used for adversarial attack detection.
The proposed approach first extracts layer activations and flattens them into activation vectors. For each selected layer, a layer-wise SVD projection basis is obtained from the corresponding layer transformation. For convolutional layers, this transformation is represented through an equivalent operator form, while linear layers are handled through their weight transformation. Sample activation vectors are then projected onto the corresponding SVD basis to obtain corevectors. Based on these projected representations, a tail-energy statistic is introduced to measure the energy of a sample in the lower singular-value components of the SVD basis. The central hypothesis is that adversarial examples may follow abnormal internal trajectories and therefore produce distinctive energy patterns in selected intermediate layers.
The method is evaluated using standard and adversarially trained Wide Residual Networks under an adversarial robustness framework involving AutoAttack components. The results show that intermediate representations contain useful statistical information for detecting certain adversarial attacks, particularly APGD-based attacks. However, detection remains more difficult for FAB and Square Attack, whose energy distributions overlap more strongly with clean samples. Overall, the study demonstrates that internal activation geometry provides a meaningful direction for adversarial detection, while highlighting the need for broader attack-aware and layer-wise analysis in future robustness research.
Tipologia del documento
Tesi di laurea
(Laurea magistrale)
Autore della tesi
Rezaei, Mohammadhossein
Relatore della tesi
Correlatore della tesi
Scuola
Corso di studio
Indirizzo
CURRICULUM ELECTRONICS FOR INTELLIGENT SYSTEMS, BIG-DATA AND INTERNET OF THINGS
Ordinamento Cds
DM270
Parole chiave
Adversarial examples, Adversarial attack detection, Intermediate activations, Singular Value Decomposition, Corevector tail energy, Wide Residual Networks, Robustness evaluation, Tail-energy analysis
Data di discussione della Tesi
20 Luglio 2026
URI
Altri metadati
Tipologia del documento
Tesi di laurea
(NON SPECIFICATO)
Autore della tesi
Rezaei, Mohammadhossein
Relatore della tesi
Correlatore della tesi
Scuola
Corso di studio
Indirizzo
CURRICULUM ELECTRONICS FOR INTELLIGENT SYSTEMS, BIG-DATA AND INTERNET OF THINGS
Ordinamento Cds
DM270
Parole chiave
Adversarial examples, Adversarial attack detection, Intermediate activations, Singular Value Decomposition, Corevector tail energy, Wide Residual Networks, Robustness evaluation, Tail-energy analysis
Data di discussione della Tesi
20 Luglio 2026
URI
Statistica sui download
Gestione del documento: