software · Zenodo (CERN European Organization for Nuclear Research)
DQShield provides an executable procedure designed to evaluate whether data quality safeguards should be kept in place for specific faults and prediction contracts. This release delivers the complete technical artefact, including the implementation, machine-readable assurance schema, frozen configurations, tests, policy sensitivity records and rerun scripts. Validation of the procedure encompasses several challenging experimental setups, such as concurrent data quality faults, regression models, temporally ordered tasks, subgroup diagnostics, and high-dimensional software vulnerability data. External datasets used for testing are referenced via their public distribution sources and checksums rather than being directly bundled.
Machine learning and analytical systems often rely on automated safeguards to catch data errors. DQShield offers an executable method to test whether these safeguards are truly necessary or effective under complex conditions, such as simultaneous data faults or evolving time-series tasks, helping practitioners maintain robust data pipelines.
This tool is relevant for software engineers, data scientists, and organisations maintaining predictive pipelines who need to audit and justify their data validation rules. Given that this artefact provides an executable implementation, schemas, and tests evaluated on realistic tasks, it represents an applied and tested resource, though commercial adoption would depend on integrating the testing framework into existing enterprise data quality operations.
AI-generated from the published abstract. Always read the original work before citing.
DQShield is an executable procedure for testing whether a data-quality safeguard should be retained for a specified fault and prediction contract. This release contains the implementation, machine-readable assurance schema, frozen configurations, run-level evidence, policy sensitivity records, tests and rerun scripts used to inspect the procedure. The accompanying validation experiments cover simultaneous data-quality faults, regression and temporally ordered tasks, subgroup diagnostics and a high-dimensional software vulnerability dataset. Third-party datasets are loaded from their public distributions or identified by source and checksum rather than redistributed.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.22542793
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.