Background: Born-digital personal archives present novel descriptive challenges that conventional archival methods are not equipped to address. Their scale, format heterogeneity, and multi-layered stratification — physical, logical, and conceptual — require a representational framework capable of capturing document structure, provenance, integrity, and relational contexts. The increasing volume of born-digital archives received by cultural heritage institutions makes the development of automated and scalable description methodologies an urgent practical priority. Methods: This paper presents a replicable five-phase pipeline, grounded in the Born-Digital Ontology (BoDi)—a formal extension of Records in Contexts Ontology (RiC-O) integrating PREMIS, PROV-O, and LRMoo — that transforms born-digital personal archive directories into RDF knowledge bases. The phases are: (1) file system ingestion, comprising SHA-256 integrity verification, parallel multi-tool metadata extraction, and triplestore loading; (2) SPARQL-based validation and rule-based enrichment, including cross-tool metadata alignment, document classification, and duplicate detection; (3) hybrid enrichment combining expert bibliographic curation with automated LLM-based description generation; (4) chainof- custody reconstruction; and (5) selective entity-level redaction for privacy-aware dissemination. Results: Evaluated on the Valerio Evangelisti Archive — 78,211 files and 11,135 folders totaling 2.1 TB across three storage media — the pipeline generated 61,154,322 RDF triples across 13 named graphs with minimal manual intervention. The resulting knowledge base supports querying across structural, typological, temporal, and relational dimensions, enabling analyses — from duplicate detection across storage media to correlation of archival activity with editorial output — that conventional finding aids cannot support. Conclusions: Born-digital archives present a double challenge: their fragility raises preservation concerns, while their scale makes manual description almost impossible. The pipeline addresses both by inverting the framing: the properties that make born-digital materials resistant to conventional methods become, within a formal framework, sources of descriptive richness without equivalent in the analog domain. Ontological modeling combined with automated processing and human-in-the-loop validation offers a viable and empirically demonstrated path toward semantic description of born-digital cultural heritage at scale.
Giagnolini, L., Bonora, P. (2026). From File Systems to Knowledge Graphs: A Multi-Phase Workflow for the Description of Born-Digital Literary Archives. IOS Press.
From File Systems to Knowledge Graphs: A Multi-Phase Workflow for the Description of Born-Digital Literary Archives
Giagnolini Lucia
;Bonora Paolo
2026
Abstract
Background: Born-digital personal archives present novel descriptive challenges that conventional archival methods are not equipped to address. Their scale, format heterogeneity, and multi-layered stratification — physical, logical, and conceptual — require a representational framework capable of capturing document structure, provenance, integrity, and relational contexts. The increasing volume of born-digital archives received by cultural heritage institutions makes the development of automated and scalable description methodologies an urgent practical priority. Methods: This paper presents a replicable five-phase pipeline, grounded in the Born-Digital Ontology (BoDi)—a formal extension of Records in Contexts Ontology (RiC-O) integrating PREMIS, PROV-O, and LRMoo — that transforms born-digital personal archive directories into RDF knowledge bases. The phases are: (1) file system ingestion, comprising SHA-256 integrity verification, parallel multi-tool metadata extraction, and triplestore loading; (2) SPARQL-based validation and rule-based enrichment, including cross-tool metadata alignment, document classification, and duplicate detection; (3) hybrid enrichment combining expert bibliographic curation with automated LLM-based description generation; (4) chainof- custody reconstruction; and (5) selective entity-level redaction for privacy-aware dissemination. Results: Evaluated on the Valerio Evangelisti Archive — 78,211 files and 11,135 folders totaling 2.1 TB across three storage media — the pipeline generated 61,154,322 RDF triples across 13 named graphs with minimal manual intervention. The resulting knowledge base supports querying across structural, typological, temporal, and relational dimensions, enabling analyses — from duplicate detection across storage media to correlation of archival activity with editorial output — that conventional finding aids cannot support. Conclusions: Born-digital archives present a double challenge: their fragility raises preservation concerns, while their scale makes manual description almost impossible. The pipeline addresses both by inverting the framing: the properties that make born-digital materials resistant to conventional methods become, within a formal framework, sources of descriptive richness without equivalent in the analog domain. Ontological modeling combined with automated processing and human-in-the-loop validation offers a viable and empirically demonstrated path toward semantic description of born-digital cultural heritage at scale.| File | Dimensione | Formato | |
|---|---|---|---|
|
SSW-63-SSW260020.pdf
accesso aperto
Descrizione: Contributo in Atti di Convegno
Tipo:
Versione (PDF) editoriale / Version Of Record
Licenza:
Licenza per Accesso Aperto. Creative Commons Attribuzione (CCBY)
Dimensione
389.92 kB
Formato
Adobe PDF
|
389.92 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



