Evaluating agentic AI systems for urban knowledge tasks: a framework based on embeddings and LLM-as-a-judge with a case study on UrbIA

Hasani, Reihaneh (2026) Evaluating agentic AI systems for urban knowledge tasks: a framework based on embeddings and LLM-as-a-judge with a case study on UrbIA. [Laurea magistrale], Università di Bologna, Corso di Studio in Physics [LM-DM270], Documento ad accesso riservato.
Documenti full-text disponibili:
[thumbnail of Thesis] Documento PDF (Thesis)
Full-text accessibile solo agli utenti istituzionali dell'Ateneo
Disponibile con Licenza: Salvo eventuali più ampie autorizzazioni dell'autore, la tesi può essere liberamente consultata e può essere effettuato il salvataggio e la stampa di una copia per fini strettamente personali di studio, di ricerca e di insegnamento, con espresso divieto di qualunque utilizzo direttamente o indirettamente commerciale. Ogni altro diritto sul materiale è riservato

Download (1MB) | Contatta l'autore

Abstract

Agentic AI systems that autonomously retrieve and reason over urban open data represent a promising direction for making city information accessible to non-expert users. However, their evaluation remains challenging: traditional reference-based metrics are poorly suited to long-form, factual responses generated by large language models (LLMs), and no established benchmark exists for the specific domain of urban informatics. This thesis introduces a reproducible evaluation framework centred on LLM-as-a-Judge and applies it to Urbia, a multi-agent system built on LangGraph framework for querying the Comune di Bologna's open data portal in natural language. The core methodological contribution is a dual-pass judging protocol in which a separate LLM evaluator (Gemini~3~Flash at zero temperature) scores each candidate response against a reference answer according to a five-criterion weighted rubric covering factual accuracy, completeness, relevance, clarity, and faithfulness to the reference. An alternative hybrid metric combining cosine-similarity-based embedding scoring with the judge score was investigated but ultimately rejected. The evaluation benchmark comprises 100 question--answer pairs grounded in Bologna's open datasets and stratified across three difficulty levels. Four current LLM backbones were compared under identical conditions via the OpenRouter API: MiniMax~M2.5, GPT~5.2, Kimi~K2.6, and Gemini~3.1~Pro. MiniMax~M2.5 achieved the highest mean judge score (82.12 out of 100), followed by GPT~5.2 (78.78), Kimi~K2.6 (75.79), and Gemini~3.1~Pro (70.83). Dual-pass agreement statistics confirmed that positional bias is negligible on this benchmark. A cost-quality analysis showed that MiniMax~M2.5 delivers both the highest accuracy and a competitive price-to-performance ratio, while Gemini~3.1~Pro offers the fastest median response time (90~s per question) at the cost of the lowest accuracy.

Abstract
Tipologia del documento
Tesi di laurea (Laurea magistrale)
Autore della tesi
Hasani, Reihaneh
Relatore della tesi
Correlatore della tesi
Scuola
Corso di studio
Indirizzo
Applied Physics
Ordinamento Cds
DM270
Parole chiave
Data-Driven Modeling,Data Analysis,Statistical Analysis,Experimental Evaluation,Benchmarking,LLM,Multi-Agent Systems,Urban Open Data,city complex science
Data di discussione della Tesi
23 Luglio 2026
URI

Altri metadati

Statistica sui download

Gestione del documento: Visualizza il documento

^