MARATTO

software · Zenodo (CERN European Organization for Nuclear Research)

regimebench: Boundary Pinning in Cluster-Count Selection -- A Controlled Benchmark

2026Open accessMohamed I University

Abstract

This deposit provides the replication artifact and benchmark suite for evaluating cluster-count selection criteria on financial and heavy-tailed time series. Overview The benchmark treats persistent, heavy-tailed volatility processes as a stress case for cluster-count selection. It evaluates standard and specialized model-selection and distance-based heuristics across synthetic ground-truth experiments, null controls, positive controls, and empirical financial data spanning a century of US equity returns. Benchmark Scope Evaluated Criteria: Silhouette, Davies–Bouldin, Calinski–Harabasz, AIC, BIC, $BIC(N_{\text{eff}})$, ICL, the Gap Statistic, and Prediction Strength. Control Regimes: Four null controls (including unstructured noise and non-clustering persistent volatility). Two positive controls: one evaluated directly on the latent state variable, and one executing identical criterion implementations on well-separated, i.i.d. Gaussian clusters. Empirical Data: 100 years of US daily market returns. Reproducibility Pipeline Every table, figure, and quantitative claim in the accompanying manuscript regenerates from a single orchestration script: Entry Point: run_all.py executes the full 17-step end-to-end pipeline. Automated Verification: The final step (tools/verify_manuscript.py) parses every numerical value reported in the LaTeX source and asserts exact equivalence against the generated data artifacts, throwing an error on drift. Integrity Self-Test: Run tools/verify_manuscript.py --self-test to verify that the verifier correctly catches and flags deliberately injected numerical corruptions. Data & Provenance Synthetic Data: All synthetic time series regenerate bit-for-bit using deterministic random seeds recorded in experiments/config/. Third-Party Data: To respect source redistributions, raw third-party payloads (Ken French Data Library, FRED, Binance) are fetched directly on first run via dedicated connectors. Integrity Audit: data/provenance.json logs the source URL, retrieval timestamp, row count, target date range, and SHA-256 hash for every external file to enable byte-for-byte verification.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.21981328

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.