SLOs and DORA
A platform that wants to be taken seriously needs explicit targets, error budgets, and the discipline to stop shipping features when the budget is exhausted.
Targets
Section titled “Targets”| SLI | Target | Window | Error budget |
| --- | --- | --- | --- |
| User-facing availability (/ returns 200 + WS handshake succeeds) | 99.9% | 30 days | 43 min/month downtime |
| Order-ack latency p99 (submit to orderAck on WS) | < 300 ms | 7 days | 1% of orders may exceed |
| WS-feed continuity (no marketUpdate gaps > 2 s) | 99.5% | 24 hours | ~7 min of feed outage / day |
| Deploy success rate | > 95% | rolling 30 deploys | One bad deploy in 20 |
| MTTR (page-able alert fires, all green again) | < 30 min | per incident | (none) |
| Change failure rate | < 15% | 30 days | 3 in 20 deploys may cause incident |
| RPO (data loss tolerated on disaster) | < 5 min | (none) | (none) |
| RTO (time to recover from total disaster) | < 30 min | (none) | (none) |
DORA metrics
Section titled “DORA metrics”| DORA metric | Target | Status | | --- | --- | --- | | Deployment frequency | Multiple per day | yes | | Lead time for changes | < 1 hour PR to prod | yes | | Change failure rate | < 15% | no, currently ~40% | | Mean time to recovery | < 1 hour | no, currently unmeasured, often hours |
The first two are "elite" by DORA's classification. The last two are "low performer". The gap is operations, not engineering velocity.