Cultural institutions interested in the art market and the history of collecting increasingly face complex challenges in digitising, analysing, and disseminating auction catalogues. While digitisation initiatives exist, systematic evaluations of workflows for digitisation, transcription, segmentation, and review remain limited, making generalisation difficult. The production of fully interoperable Linked Open Data is rare, with semantic enrichment often limited to small data subsets. The Federico Zeri Foundation has launched a project to digitise 1,900 auction catalogues spanning 1869-1929 across six countries. This effort employs an iterative workflow combining automated rule-based and LLM methods with human validation to create semantically enriched representations. Acknowledging that digitisation choices influence historical interpretation, an incremental data-modelling strategy is adopted to accommodate variable transcription and segmentation quality. This article details such workflow and provides a preliminary evaluation of its effectiveness.
Daquino, M., Mambelli, F., Pasqual, V., Rossetti, V., Tomasi, F. (2026). Tracing the art market: a hybrid approach to digitise historical auction catalogues. CEUR-WS.
Tracing the art market: a hybrid approach to digitise historical auction catalogues
Daquino Marilena;Pasqual Valentina;Rossetti Valentina;Tomasi Francesca
2026
Abstract
Cultural institutions interested in the art market and the history of collecting increasingly face complex challenges in digitising, analysing, and disseminating auction catalogues. While digitisation initiatives exist, systematic evaluations of workflows for digitisation, transcription, segmentation, and review remain limited, making generalisation difficult. The production of fully interoperable Linked Open Data is rare, with semantic enrichment often limited to small data subsets. The Federico Zeri Foundation has launched a project to digitise 1,900 auction catalogues spanning 1869-1929 across six countries. This effort employs an iterative workflow combining automated rule-based and LLM methods with human validation to create semantically enriched representations. Acknowledging that digitisation choices influence historical interpretation, an incremental data-modelling strategy is adopted to accommodate variable transcription and segmentation quality. This article details such workflow and provides a preliminary evaluation of its effectiveness.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



