A Multi-Agent LLM Pipeline for Argument Mining on Legal Judgments

Banihashemi, Safoura (2026) A Multi-Agent LLM Pipeline for Argument Mining on Legal Judgments. [Laurea magistrale], Università di Bologna, Corso di Studio in Artificial intelligence [LM-DM270]
Documenti full-text disponibili:
[thumbnail of Thesis] Documento PDF (Thesis)
Disponibile con Licenza: Creative Commons: Attribuzione - Non commerciale - Non opere derivate 4.0 (CC BY-NC-ND 4.0)

Download (6MB)

Abstract

Legal Argumentation Mining (LAM) aims to uncover the argumentative structure in legal documents. It identifies which sentences present arguments, their roles, the types of reasoning they use, and how they support or challenge each other. This process naturally breaks down into several related sub-tasks. However, most research focuses on each sub-task in isolation, and it remains unclear how these sub-tasks work together within a complete system. This thesis presents the design, implementation, and evaluation of a five-agent LLM pipeline for LAM, using the Demosthenes corpus, an annotated dataset of decisions on fiscal State aid by the Court of Justice of the European Union. The pipeline uses a LangGraph state machine to coordinate agents that handle the different sub tasks: detection of argumentative components, their classification in premises or conclusions, classification of factual and legal premises, identification of argumentation schemes, and directed relation identification. The main methodological contribution is an evaluation framework to measure the performance of the system under three different settings, to quantify the intrinsic capabilities of individual agents and the impact of error propagation. First, each agent is executed in an isolated setting and evaluated on the golden data. Then, the full pipeline is evaluated both in an upstream-conditioned setting, to evaluate the agents without penalizing them for mistakes made by their predecessors, and in an end-to-end setting to evaluate the performance of the pipeline as a whole, without distinguishing the contribution of single agents. The experiments compare a commercial model (Gemini) with an open-weight model (Gemma), and show the importance of evaluating modular LLM systems on their overall performance, not just on individual tasks.

Abstract
Tipologia del documento
Tesi di laurea (Laurea magistrale)
Autore della tesi
Banihashemi, Safoura
Relatore della tesi
Correlatore della tesi
Scuola
Corso di studio
Ordinamento Cds
DM270
Parole chiave
Legal Argument Mining, Large Language Models, Multi-Agent Systems, Prompt Engineering, Error Propagation, Argumentation Schemes, Court of Justice of the European Union
Data di discussione della Tesi
21 Luglio 2026
URI

Altri metadati

Statistica sui download

Gestione del documento: Visualizza il documento

^