MARATTO

peer review · Zenodo (CERN European Organization for Nuclear Research)

Ethical Guardrail Bandit Pruning (EGBP): A Fairness-Constrained Reinforcement Learning System for Equitable Governance of Distributed Bioenergy Grids

2026Open accessNorth-West University

Abstract

131,400 node-hour observations across 15 bioenergy grid nodes (6 anaerobic digesters, 5 gasifiers, 4 CHP units) over 365 days at hourly resolution. Parameters calibrated to Faaij (2006) and Atashbar et al. (2016). Generated for: Ethical Guardrail Bandit Pruning (EGBP). This repository supports the manuscript "Ethical Guardrail Bandit Pruning (EGBP): A Framework for Sustainable and Equitable AI-Driven Resource Allocation in Bioenergy Grids" (under review). EGBP treats computational efficiency, energy sustainability, and distributional equity as simultaneously binding optimisation constraints rather than sequentially addressed objectives. The central theoretical contribution demonstrates that embedding ethical guardrails directly in the optimisation objective, rather than evaluating them post-hoc preserves the asymptotic regret properties of Thompson Sampling bandit learning while guaranteeing convergence to an ethically-admissible stationary point. The framework integrates four tightly coupled mechanisms: (i) a hybrid bandit-gradient importance estimator combining offline Transformer pre-training with online Thompson Sampling; (ii) an ethical guardrail buffer enforcing differentiable penalties on energy overconsumption and distributional inequity through a Gini-coefficient regulariser applied to physical energy budget allocations; (iii) a cost-weighted magnitude pruning operator coupling gradient sparsity to real-time biomass conversion telemetry; and (iv) a guardrail-filtered federated averaging scheme with analytically bounded exclusion fraction. Theoretical properties are established through three theorems, three propositions, and two corollaries covering energy guardrail self-correction, Gini penalty convexity, Thompson Sampling regret preservation, EGBP convergence, and federated fairness monotonicity. Simulation on a 15-node bioenergy grid over 50,000 training steps demonstrates 38% reduction in energy consumption and 40% improvement in distributional fairness relative to three competitive baselines. This repository contains derived simulation outputs and pipeline code for reproducibility.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.22046887

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.