Accurate extraction of argumentative information from legal documents allows transforming legal documents into structured representations that can support downstream applications, such as information retrieval, document analysis, and legal decision support. Legal Argument Mining addresses the identification and classification of argumentative structures and their relations in legal texts. Such a task can be decomposed into multiple interdependent reasoning steps. In this work, we investigate the implementation of an argument mining pipeline using LLM agents, each one dedicated to a separate task. We evaluate our system under three different settings that progressively expose the impact of error propagation, ranging from independent evaluation of each agent to a realistic end-to-end scenario. We perform an experimental evaluation on the Demosthenes corpus, comparing an open-weight model against a commercial one. Results show that cascading errors substantially degrade downstream performance, making the first stages of the pipeline a significant bottleneck. Our findings highlight the importance of evaluating modular LLM systems in their end-to-end behavior in addition to individual task performance.
Banihashemi, S., Grundler, G., Galassi, A. (2026). A Multi-Agent LLM Pipeline for Legal Argument Mining [10.1145/3820755.3833404].
A Multi-Agent LLM Pipeline for Legal Argument Mining
Giulia Grundler
;Andrea Galassi
2026
Abstract
Accurate extraction of argumentative information from legal documents allows transforming legal documents into structured representations that can support downstream applications, such as information retrieval, document analysis, and legal decision support. Legal Argument Mining addresses the identification and classification of argumentative structures and their relations in legal texts. Such a task can be decomposed into multiple interdependent reasoning steps. In this work, we investigate the implementation of an argument mining pipeline using LLM agents, each one dedicated to a separate task. We evaluate our system under three different settings that progressively expose the impact of error propagation, ranging from independent evaluation of each agent to a realistic end-to-end scenario. We perform an experimental evaluation on the Demosthenes corpus, comparing an open-weight model against a commercial one. Results show that cascading errors substantially degrade downstream performance, making the first stages of the pipeline a significant bottleneck. Our findings highlight the importance of evaluating modular LLM systems in their end-to-end behavior in addition to individual task performance.| File | Dimensione | Formato | |
|---|---|---|---|
|
3820755.3833404.pdf
accesso aperto
Tipo:
Versione (PDF) editoriale / Version Of Record
Licenza:
Licenza per Accesso Aperto. Creative Commons Attribuzione (CCBY)
Dimensione
377.15 kB
Formato
Adobe PDF
|
377.15 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



