Automated sensitivity classification could reduce manual review of government documents, but benchmark scores are unsafe when taxonomies, retrieval indexes, and prediction provenance are not auditable. This paper uses Mikyas, a prototype spanning six large language models, four prompting conditions, and 10,106 documents, as an artifact case study for an audit-first evaluation framework—not a comparative performance study. The frozen 242,544-row matrix contains 428 live observations, 684 traceably reused observations, 235,532 deterministic projections, and 5,900 unverified Run-12 matches whose independence cannot be demonstrated. A project log records top-ranked scores of at least 0.99 for 98 of 100 uniformly sampled self-queries against a non-isolated retrieval index. We make provenance, retrieval isolation, manifest binding, annotation traceability, and human approval conditions for metric eligibility, and distinguish cloud, sovereign-connected, and air-gapped evaluation profiles. Mikyas therefore provides an audit-first research framework, not evidence of model superiority or automated regulatory compliance.
Al Darmaki, A., Martino, L., Yeob Yeun, C., Al Hammadi, Y. (2026). Mikyas: Audit-First Evaluation of Large Language Models for UAE Government Data Classification.
Mikyas: Audit-First Evaluation of Large Language Models for UAE Government Data Classification
Luigi Martino;
2026
Abstract
Automated sensitivity classification could reduce manual review of government documents, but benchmark scores are unsafe when taxonomies, retrieval indexes, and prediction provenance are not auditable. This paper uses Mikyas, a prototype spanning six large language models, four prompting conditions, and 10,106 documents, as an artifact case study for an audit-first evaluation framework—not a comparative performance study. The frozen 242,544-row matrix contains 428 live observations, 684 traceably reused observations, 235,532 deterministic projections, and 5,900 unverified Run-12 matches whose independence cannot be demonstrated. A project log records top-ranked scores of at least 0.99 for 98 of 100 uniformly sampled self-queries against a non-isolated retrieval index. We make provenance, retrieval isolation, manifest binding, annotation traceability, and human approval conditions for metric eligibility, and distinguish cloud, sovereign-connected, and air-gapped evaluation profiles. Mikyas therefore provides an audit-first research framework, not evidence of model superiority or automated regulatory compliance.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



