Background: Born-digital personal archives present novel descriptive challenges that conventional archival methods are not equipped to address. Their scale, format heterogeneity, and multi-layered stratification — physical, logical, and conceptual — require a representational framework capable of capturing document structure, provenance, integrity, and relational contexts. The increasing volume of born-digital archives received by cultural heritage institutions makes the development of automated and scalable description methodologies an urgent practical priority. Methods: This paper presents a replicable five-phase pipeline, grounded in the Born-Digital Ontology (BoDi)—a formal extension of Records in Contexts Ontology (RiC-O) integrating PREMIS, PROV-O, and LRMoo — that transforms born-digital personal archive directories into RDF knowledge bases. The phases are: (1) file system ingestion, comprising SHA-256 integrity verification, parallel multi-tool metadata extraction, and triplestore loading; (2) SPARQL-based validation and rule-based enrichment, including cross-tool metadata alignment, document classification, and duplicate detection; (3) hybrid enrichment combining expert bibliographic curation with automated LLM-based description generation; (4) chainof- custody reconstruction; and (5) selective entity-level redaction for privacy-aware dissemination. Results: Evaluated on the Valerio Evangelisti Archive — 78,211 files and 11,135 folders totaling 2.1 TB across three storage media — the pipeline generated 61,154,322 RDF triples across 13 named graphs with minimal manual intervention. The resulting knowledge base supports querying across structural, typological, temporal, and relational dimensions, enabling analyses — from duplicate detection across storage media to correlation of archival activity with editorial output — that conventional finding aids cannot support. Conclusions: Born-digital archives present a double challenge: their fragility raises preservation concerns, while their scale makes manual description almost impossible. The pipeline addresses both by inverting the framing: the properties that make born-digital materials resistant to conventional methods become, within a formal framework, sources of descriptive richness without equivalent in the analog domain. Ontological modeling combined with automated processing and human-in-the-loop validation offers a viable and empirically demonstrated path toward semantic description of born-digital cultural heritage at scale.

Giagnolini, L., Bonora, P. (2026). From File Systems to Knowledge Graphs: A Multi-Phase Workflow for the Description of Born-Digital Literary Archives. IOS Press.

From File Systems to Knowledge Graphs: A Multi-Phase Workflow for the Description of Born-Digital Literary Archives

Giagnolini Lucia
;
Bonora Paolo
2026

Abstract

Background: Born-digital personal archives present novel descriptive challenges that conventional archival methods are not equipped to address. Their scale, format heterogeneity, and multi-layered stratification — physical, logical, and conceptual — require a representational framework capable of capturing document structure, provenance, integrity, and relational contexts. The increasing volume of born-digital archives received by cultural heritage institutions makes the development of automated and scalable description methodologies an urgent practical priority. Methods: This paper presents a replicable five-phase pipeline, grounded in the Born-Digital Ontology (BoDi)—a formal extension of Records in Contexts Ontology (RiC-O) integrating PREMIS, PROV-O, and LRMoo — that transforms born-digital personal archive directories into RDF knowledge bases. The phases are: (1) file system ingestion, comprising SHA-256 integrity verification, parallel multi-tool metadata extraction, and triplestore loading; (2) SPARQL-based validation and rule-based enrichment, including cross-tool metadata alignment, document classification, and duplicate detection; (3) hybrid enrichment combining expert bibliographic curation with automated LLM-based description generation; (4) chainof- custody reconstruction; and (5) selective entity-level redaction for privacy-aware dissemination. Results: Evaluated on the Valerio Evangelisti Archive — 78,211 files and 11,135 folders totaling 2.1 TB across three storage media — the pipeline generated 61,154,322 RDF triples across 13 named graphs with minimal manual intervention. The resulting knowledge base supports querying across structural, typological, temporal, and relational dimensions, enabling analyses — from duplicate detection across storage media to correlation of archival activity with editorial output — that conventional finding aids cannot support. Conclusions: Born-digital archives present a double challenge: their fragility raises preservation concerns, while their scale makes manual description almost impossible. The pipeline addresses both by inverting the framing: the properties that make born-digital materials resistant to conventional methods become, within a formal framework, sources of descriptive richness without equivalent in the analog domain. Ontological modeling combined with automated processing and human-in-the-loop validation offers a viable and empirically demonstrated path toward semantic description of born-digital cultural heritage at scale.
2026
Proceedings of the 22nd International Conference on Semantic Systems
229
245
Giagnolini, L., Bonora, P. (2026). From File Systems to Knowledge Graphs: A Multi-Phase Workflow for the Description of Born-Digital Literary Archives. IOS Press.
Giagnolini, Lucia; Bonora, Paolo
File in questo prodotto:
File Dimensione Formato  
SSW-63-SSW260020.pdf

accesso aperto

Descrizione: Contributo in Atti di Convegno
Tipo: Versione (PDF) editoriale / Version Of Record
Licenza: Licenza per Accesso Aperto. Creative Commons Attribuzione (CCBY)
Dimensione 389.92 kB
Formato Adobe PDF
389.92 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11585/1083390
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact