Jamshidi, Mohammad
(2026)
An Investigation on Generative-AI and RAG Support to Question Answering on Scientific Papers.
[Laurea magistrale], Università di Bologna, Corso di Studio in
Digital transformation management [LM-DM270] - Cesena, Documento full-text non disponibile
Il full-text non è disponibile per scelta dell'autore.
(
Contatta l'autore)
Abstract
Retrieval-Augmented Generation (RAG) has become the prominent architectural framework in order to automate the answer and assessment of open-ended questions (OEQs) in educational environments. However, this process’s effectiveness at each stage of the retrieval process is influenced by various architectural design decisions, such as document segmentation, choice of embedding models, and similarity-based retrieval. This thesis aims to empirically investigate in a controlled procedure, based on the retrieval pipeline stages, the notion of retrieval quality in the context of OEQ-oriented RAG systems.
The corpus has been carefully selected and composed of 22 scientific documents and 19 OEQs, with ground-truth relevance labels for each document-question pair, annotated manually. Two chunking strategies were examined: fixed-length chunking and structure-aware chunking, combined with four selected Sentencetransformer embedding models, namely the all-mpnet-base-v2, all-MiniLM-L12 -v2, all-MiniLM-L6-v2, and all-DistilRoBERTa-v1. Precision, recall, and F1- score metrics have been used for Retrieval-performance assessment at a osinesimilarity, similarity of text-question, threshold of 0.5.. The results show that structure-aware chunking is superior to fixed-length chunking. Among the experimented embedding models, all-mpnet-base-v2 performed in a superior way regarding the retrieval results. The findings further highlight that question formulation quality and the presence of bibliographic noise
are significant sources of false positives, underscoring the mportance of an adequate pre-processing stage and question designs, especially in automated assessment pipelines.
Abstract
Retrieval-Augmented Generation (RAG) has become the prominent architectural framework in order to automate the answer and assessment of open-ended questions (OEQs) in educational environments. However, this process’s effectiveness at each stage of the retrieval process is influenced by various architectural design decisions, such as document segmentation, choice of embedding models, and similarity-based retrieval. This thesis aims to empirically investigate in a controlled procedure, based on the retrieval pipeline stages, the notion of retrieval quality in the context of OEQ-oriented RAG systems.
The corpus has been carefully selected and composed of 22 scientific documents and 19 OEQs, with ground-truth relevance labels for each document-question pair, annotated manually. Two chunking strategies were examined: fixed-length chunking and structure-aware chunking, combined with four selected Sentencetransformer embedding models, namely the all-mpnet-base-v2, all-MiniLM-L12 -v2, all-MiniLM-L6-v2, and all-DistilRoBERTa-v1. Precision, recall, and F1- score metrics have been used for Retrieval-performance assessment at a osinesimilarity, similarity of text-question, threshold of 0.5.. The results show that structure-aware chunking is superior to fixed-length chunking. Among the experimented embedding models, all-mpnet-base-v2 performed in a superior way regarding the retrieval results. The findings further highlight that question formulation quality and the presence of bibliographic noise
are significant sources of false positives, underscoring the mportance of an adequate pre-processing stage and question designs, especially in automated assessment pipelines.
Tipologia del documento
Tesi di laurea
(Laurea magistrale)
Autore della tesi
Jamshidi, Mohammad
Relatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
Retrieval,Augmented,Generation,Embedding,Retrieval, Process,Chunking,Open,Ended,Question
Data di discussione della Tesi
15 Luglio 2026
URI
Altri metadati
Tipologia del documento
Tesi di laurea
(NON SPECIFICATO)
Autore della tesi
Jamshidi, Mohammad
Relatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
Retrieval,Augmented,Generation,Embedding,Retrieval, Process,Chunking,Open,Ended,Question
Data di discussione della Tesi
15 Luglio 2026
URI
Gestione del documento: