MARATTO

software · Zenodo (CERN European Organization for Nuclear Research)

Reproducibility Artifact for "Benchmark Success Does Not Guarantee Operational Assurance: A Prospective Multi-Environment Transport Audit of Industrial Anomaly Detectors"

2026Open accessLead City University

In plain language

This release provides a comprehensive reproducibility package designed to support a multi-environment transport audit of industrial anomaly detection models. It contains the locked development benchmarks, calibration records, confirmatory score streams, event ledgers, and diagnostic analyses used to scrutinise whether model success on standard benchmarks translates to reliable operational assurance. The dataset preserves evaluation materials covering held-out environments, including HAI 22.04, Sherlock Basic, and Sherlock Semiurban, while tracking specific decisions such as the PCAResidual method. It incorporates verification reports, source code, and correction manifests, specifically noting where earlier event-derived outputs for Sherlock Basic have been superseded by a corrected evaluation of eighteen attacks. Access to the underlying raw industrial datasets must be obtained separately from their primary sources.

Key takeaways

  • The artifact provides full calibration logs, verification reports, and source code auditing the real-world operational assurance of industrial anomaly detectors.
  • It maintains evaluation records across multiple independent test environments, notably HAI 22.04 and two distinct Sherlock system settings.
  • The archive preserves a specific decision regarding the PCAResidual detector alongside diagnostic post-hoc analyses.
  • Evaluation updates are explicitly tracked, designating earlier Sherlock Basic event outputs as superseded in favour of an official eighteen-attack assessment.
  • Raw data from industrial testbeds such as SWaT, HAI, and Sherlock are excluded to respect third-party licensing.

Why it matters

Industrial systems rely on anomaly detectors to prevent failures and detect cyber attacks, but high scores in laboratory tests do not guarantee reliable performance in operation. Providing open verification tools and data audits helps researchers and operators confirm whether safety-critical monitoring systems function properly across different industrial environments.

Commercialisation angle

The materials can be used by industrial automation engineers, security auditors, and system integrators seeking to independently verify or test anomaly detection tools prior to plant deployment. Because this artifact serves as an analytical audit and benchmarking suite rather than a standalone commercial product, it functions at the validation and testing stage of the technology development pipeline.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

This reproducibility artifact contains the frozen development benchmark lock, confirmatory score streams and calibration records, event ledgers, correction manifests, post-hoc diagnostic analyses, source code, and verification reports supporting a multi-environment audit of industrial anomaly detectors. The artifact preserves the development-supported PCAResidual decision and the held-out HAI 22.04, Sherlock Basic, and Sherlock Semiurban evidence used to evaluate deployment-facing assurance screening. Raw third-party SWaT, HAI, and Sherlock datasets are not redistributed; users must obtain them from their official sources under the applicable licences. The Sherlock Basic directory preserves both the original frozen-score evidence and the corrected official 18-attack event evaluation, with the original event-derived outputs explicitly marked as superseded for scientific inference.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.22067730

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.