article · Zenodo (CERN European Organization for Nuclear Research)
This archive provides a complete computational and reproducibility package for benchmarking machine learning models applied to biomass gasification prediction. It contains empirical records from six direct QRO-401 analyser measurements alongside a 98-row dataset of derived performance scenarios. The materials include fixed random seeds, frozen model specifications, repeated cross-validation generators, and automated detectors for data lineage and target proximity. It also supplies leave-configuration-out transport tests and verification outputs. Within the benchmarking structure, feature sets designated as Tier A and Tier B serve as legitimate interpolation benchmarks. In contrast, the Tier C feature set contains formula-proximal energy production metrics and is retained solely as a diagnostic test for data leakage and formula recovery, rather than as an operational prediction tool.
Machine learning models in energy research require rigorous testing to ensure their predictions reflect real physical dynamics rather than statistical leakage. By releasing complete data lineage, verification scripts, and clear tier distinctions, this resource enables researchers to independently audit, test, and replicate predictive methods on small experimental energy datasets.
The archive offers an early-stage methodological resource for researchers and software developers evaluating predictive machine learning models in biomass gasification. However, because the repository is designed strictly for computational verification, auditing, and benchmarking rather than field deployment, the abstract does not indicate a direct commercialisation pathway or operational readiness level.
AI-generated from the published abstract. Always read the original work before citing.
Executable Online Resource and reproducibility archive accompanying the manuscript “Interpretable Benchmarking of Machine Learning Models for Small Experimental Energy Systems: Balancing Accuracy, Complexity and Physical Meaning in Biomass Gasification Prediction.” The archive contains the six direct QRO-401 analyzer records retained as empirical provenance anchors and the 98-row workbook-derived performance scenario dataset used for methodological benchmarking. The 98 scenarios are derived analytical records and must not be interpreted as 98 independent physical gasifier experiments. The repository provides the fixed analysis seed (20260901), frozen model specifications, complete repeated cross-validation generator, automated data-lineage and target-proximity detector, leave-configuration-out transport tests, full fold- and repeat-level computational results, manuscript-result verification outputs, machine-readable Online Resource tables, supplementary information, manuscript-aligned figures, pinned software environment files, repository manifest, and SHA-256 checksums. Tier A and Tier B constitute the legitimate interpolation benchmarking feature sets. Tier C includes formula-proximal energy production and is retained strictly as a leakage and formula-recovery diagnostic rather than as a deployable prediction benchmark. The archive is intended to support independent inspection, computational reproduction, provenance auditing, and verification of the results reported in the associated manuscript.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.22396885
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.