software · Zenodo (CERN European Organization for Nuclear Research)
DQShield provides an executable procedure designed to test whether a data-quality safeguard ought to be kept for a given fault and prediction contract. This artefact release includes the full implementation, a machine-readable assurance schema, frozen configurations, test scripts, and run-level evidence alongside policy sensitivity records. The procedure was evaluated across experiments involving simultaneous data-quality faults, temporally ordered tasks, regression problems, and subgroup diagnostics. It was also assessed using a high-dimensional software vulnerability dataset, relying on external public distributions referenced by source and checksum.
Data errors can degrade predictive models, making safeguards essential for reliable operation. This work provides a concrete, reproducible method and assurance schema for assessing whether specific quality defences truly remain beneficial under diverse fault conditions and across different types of prediction tasks.
This procedure could be used by software engineers and machine learning practitioners who need to validate data-quality defences before deploying models into production. The artefact appears applied and tested within experimental settings, including high-dimensional vulnerability datasets, though further adaptation would be required to integrate it into commercial software development pipelines.
AI-generated from the published abstract. Always read the original work before citing.
DQShield is an executable procedure for testing whether a data-quality safeguard should be retained for a specified fault and prediction contract. This release contains the implementation, machine-readable assurance schema, frozen configurations, run-level evidence, policy sensitivity records, tests and rerun scripts used to inspect the procedure. The accompanying validation experiments cover simultaneous data-quality faults, regression and temporally ordered tasks, subgroup diagnostics and a high-dimensional software vulnerability dataset. Third-party datasets are loaded from their public distributions or identified by source and checksum rather than redistributed.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.22542792
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.