Accurate extraction of argumentative information from legal documents allows transforming legal documents into structured representations that can support downstream applications, such as information retrieval, document analysis, and legal decision support. Legal Argument Mining addresses the identification and classification of argumentative structures and their relations in legal texts. Such a task can be decomposed into multiple interdependent reasoning steps. In this work, we investigate the implementation of an argument mining pipeline using LLM agents, each one dedicated to a separate task. We evaluate our system under three different settings that progressively expose the impact of error propagation, ranging from independent evaluation of each agent to a realistic end-to-end scenario. We perform an experimental evaluation on the Demosthenes corpus, comparing an open-weight model against a commercial one. Results show that cascading errors substantially degrade downstream performance, making the first stages of the pipeline a significant bottleneck. Our findings highlight the importance of evaluating modular LLM systems in their end-to-end behavior in addition to individual task performance.
Banihashemi, S., Grundler, G., Galassi, A. (2026). A Multi-Agent LLM Pipeline for Legal Argument Mining [10.1145/3820755.3833404].
A Multi-Agent LLM Pipeline for Legal Argument Mining
Giulia Grundler
;Andrea Galassi
2026
Abstract
Accurate extraction of argumentative information from legal documents allows transforming legal documents into structured representations that can support downstream applications, such as information retrieval, document analysis, and legal decision support. Legal Argument Mining addresses the identification and classification of argumentative structures and their relations in legal texts. Such a task can be decomposed into multiple interdependent reasoning steps. In this work, we investigate the implementation of an argument mining pipeline using LLM agents, each one dedicated to a separate task. We evaluate our system under three different settings that progressively expose the impact of error propagation, ranging from independent evaluation of each agent to a realistic end-to-end scenario. We perform an experimental evaluation on the Demosthenes corpus, comparing an open-weight model against a commercial one. Results show that cascading errors substantially degrade downstream performance, making the first stages of the pipeline a significant bottleneck. Our findings highlight the importance of evaluating modular LLM systems in their end-to-end behavior in addition to individual task performance.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



