We introduce the first benchmark specifically designed to evaluate the capabilities of large language models on historical Italian. While benchmarks for modern Italian are available, historical varieties of the language remain underexplored in this context. To address this gap, we repurpose and adapt several high-quality digital resources to construct a novel benchmark. The result spans a range of tasks, including word sense disambiguation, syntactic structure prediction, intertextuality detection, named entity recognition, authorship verification, and chronological detection, with source texts ranging from the High Middle Ages to the nineteenth century. We evaluate several state-of-the-art large language models in a zero-shot setting. Our findings show that while these models exhibit promising performance, the benchmark remains far from saturated, suggesting that the capabilities of large language models on historical languages still have considerable margin for improvement. Our work offers a standardized evaluation framework to support the selection of models for annotating digital archives and to guide the development of linguistically inclusive language models.
Zhang, S., Levchenko, M., Manca, C., Sabba, F., Italia, P.M.C., Colavizza, G. (2026). Repurposing and Adapting Language Resources for a Historical Italian Large Language Model Benchmark. JOURNAL OF OPEN HUMANITIES DATA, 12, 119-137 [10.5334/johd.538].
Repurposing and Adapting Language Resources for a Historical Italian Large Language Model Benchmark
Zhang, Shibingfeng;Levchenko, Mariia;Manca, Chiara;Sabba, Fiammetta;Italia, Paola Maria Carmela;Colavizza, Giovanni
2026
Abstract
We introduce the first benchmark specifically designed to evaluate the capabilities of large language models on historical Italian. While benchmarks for modern Italian are available, historical varieties of the language remain underexplored in this context. To address this gap, we repurpose and adapt several high-quality digital resources to construct a novel benchmark. The result spans a range of tasks, including word sense disambiguation, syntactic structure prediction, intertextuality detection, named entity recognition, authorship verification, and chronological detection, with source texts ranging from the High Middle Ages to the nineteenth century. We evaluate several state-of-the-art large language models in a zero-shot setting. Our findings show that while these models exhibit promising performance, the benchmark remains far from saturated, suggesting that the capabilities of large language models on historical languages still have considerable margin for improvement. Our work offers a standardized evaluation framework to support the selection of models for annotating digital archives and to guide the development of linguistically inclusive language models.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



