software · Zenodo (CERN European Organization for Nuclear Research)
A controlled benchmark for cluster-count selection on time series, using persistentheavy-tailed volatility processes as a deliberately hard case. It benchmarks Silhouette,Davies-Bouldin, Calinski-Harabasz, AIC, BIC, BIC(N_eff), ICL, the Gap Statistic andPrediction Strength against known ground truth, four null controls, two positive controls(one on the latent state variable, one running the identical criterion code onwell-separated i.i.d. Gaussian clusters), and a century of US daily returns. Every table and figure in the accompanying article regenerates from a single entry point(run_all.py), a seventeen-step pipeline whose final step is tools/verify_manuscript.py: itparses each reported number out of the LaTeX source and asserts it against the artifactthat produced it, failing on drift. A --self-test flag confirms the verifier detectsinjected corruption. Data: synthetic series are regenerated bit-for-bit from a seed recorded inexperiments/config/. Raw third-party payloads (Ken French Data Library, FRED, Binance) arenot redistributed here; the connectors fetch them from source on first run anddata/provenance.json records each payload's source URL, retrieval timestamp, row count,date range and SHA-256, so a reader verifies byte-for-byte.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.21981329
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.