Automated sensitivity classification could reduce manual review of government documents, but benchmark scores are unsafe when taxonomies, retrieval indexes, and prediction provenance are not auditable. This paper uses Mikyas, a prototype spanning six large language models, four prompting conditions, and 10,106 documents, as an artifact case study for an audit-first evaluation framework—not a comparative performance study. The frozen 242,544-row matrix contains 428 live observations, 684 traceably reused observations, 235,532 deterministic projections, and 5,900 unverified Run-12 matches whose independence cannot be demonstrated. A project log records top-ranked scores of at least 0.99 for 98 of 100 uniformly sampled self-queries against a non-isolated retrieval index. We make provenance, retrieval isolation, manifest binding, annotation traceability, and human approval conditions for metric eligibility, and distinguish cloud, sovereign-connected, and air-gapped evaluation profiles. Mikyas therefore provides an audit-first research framework, not evidence of model superiority or automated regulatory compliance.

Al Darmaki, A., Martino, L., Yeob Yeun, C., Al Hammadi, Y. (2026). Mikyas: Audit-First Evaluation of Large Language Models for UAE Government Data Classification.

Mikyas: Audit-First Evaluation of Large Language Models for UAE Government Data Classification

Luigi Martino;
2026

Abstract

Automated sensitivity classification could reduce manual review of government documents, but benchmark scores are unsafe when taxonomies, retrieval indexes, and prediction provenance are not auditable. This paper uses Mikyas, a prototype spanning six large language models, four prompting conditions, and 10,106 documents, as an artifact case study for an audit-first evaluation framework—not a comparative performance study. The frozen 242,544-row matrix contains 428 live observations, 684 traceably reused observations, 235,532 deterministic projections, and 5,900 unverified Run-12 matches whose independence cannot be demonstrated. A project log records top-ranked scores of at least 0.99 for 98 of 100 uniformly sampled self-queries against a non-isolated retrieval index. We make provenance, retrieval isolation, manifest binding, annotation traceability, and human approval conditions for metric eligibility, and distinguish cloud, sovereign-connected, and air-gapped evaluation profiles. Mikyas therefore provides an audit-first research framework, not evidence of model superiority or automated regulatory compliance.
2026
Conference Proceedings of the IEEE-CH Cyber Humanities 2026
1
6
Al Darmaki, A., Martino, L., Yeob Yeun, C., Al Hammadi, Y. (2026). Mikyas: Audit-First Evaluation of Large Language Models for UAE Government Data Classification.
Al Darmaki, Ahmed; Martino, Luigi; Yeob Yeun, Chan; Al Hammadi, Yousef
File in questo prodotto:
Eventuali allegati, non sono esposti

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11585/1082511
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact