Limitations
Seven things this system does not solve. They are the terms on which it is being built, not a disclaimer.
Argued in full in the whitepaper at §15.
The seven#
| Limitation | What it means in practice |
|---|---|
| Cold start | Post-campaign retention is the highest-signal feature and is unavailable for pools that have never seen a campaign end — on a new chain, nearly all of them, which is exactly where durability information is most valuable |
| Aggregate-only v1 | Concentration, median tenure and tenure Gini are estimated from partial event data with materially wider error until address-level indexing lands |
| Exogenous breaches | Exploits, depegs and governance failures breach attestations without being predictable from incentive dynamics. They count. Correct for the guarantee, costly for measured accuracy |
| Reflexivity | Mitigated, not solved — see below |
| Single-attester bootstrap | At launch there is one attester. The bond is real and the challenge is permissionless, but there is no k-of-n redundancy and no diversity of method |
| Breach-definition risk | With no dispute layer, the definition carries the entire adjudication burden. A manipulable accounting method or a mis-set window has no human backstop |
| Regime coverage | Conformal guarantees hold under exchangeability, so they lapse precisely during stress — an admission that the guarantee is weakest when it is most wanted |
Reflexivity#
A widely consumed four-day horizon may cause the outflow it predicts. Agents read it, withdraw, trigger the breach, and the attestation is validated by its own publication. The forecast becomes an instruction, and the accuracy record scores it as a hit.
Three partial responses:
- Publish level, not delta. A falling horizon is directly tradeable as momentum; a stated level is less so. No trend arrow, no downgrade broadcast.
- Measure and disclose. The induced component is estimable against matched unattested controls, and it is published as a standing metric even when it undercuts headline accuracy.
- Accept the asymmetry. Reflexivity bites on short horizons and barely on long ones, which argues for conservative short-end thresholds.
How accuracy is reported#
Every metric is computed from the public attestation record, so any third party can recompute it — a metric only Cleaton can compute is a metric Cleaton can shade.
| Metric | Why it is there |
|---|---|
| Coverage rate by confidence bucket | The core calibration check |
| Sharpness | Mean horizon conditional on survival. Without it, publishing conservatively short horizons everywhere would score perfectly and be useless |
| Attestation coverage | Fraction of eligible pools attested. 99% accuracy on 12% coverage is a far weaker claim than 91% on 80% |
| Reflexive component | Estimated induced outflow versus matched controls |
| Bond ratio | Total bond over total gated exposure |
Accuracy is reported against three baselines, not in isolation: campaign expiry, a constant historical median, and an elasticity-only single-feature model. If the ensemble does not materially beat the campaign-expiry baseline, the correct action is to publish that finding. That is stated in advance so it is a commitment rather than a possible future concession.