werk 15 · reliability and response time

PerfOps.

How fast a system answers and how often it answers at all — service-level objectives set with the business, capacity modelled ahead of the peak, and a performance budget enforced in the pipeline rather than argued after release.

FamilyPerformanceCarried byAdvisoryCarried byInterimProven lines1 / 4Detailed mandates1IndustriesBetting

Werk 15 — who carries it

arms and mandates

Availability and latency are promises, and a promise with no budget attached is a hope. This werk states the promise in numbers, spends against it deliberately, and enforces it in the pipeline instead of in the post-mortem.

The offer

what is bought, and in what shape

A production estate with stated service levels, an error budget that has teeth, and response-time regressions caught before a release rather than by a customer.

Service-level objectives and error budgets

Availability and latency targets set with the business rather than inherited from a dashboard, with the escalation and release-freeze rule that gives an error budget consequence.

Performance budget in the pipeline

Response-time thresholds enforced as a build gate, load and soak tests run against a representative dataset, and the regression traced to the change that caused it.

Capacity and scalability planning

Peak-load modelling, headroom stated in numbers, and the scaling behaviour proven under failure before the event that needs it.

Observability and incident practice

Signals that answer "is it the user's experience" rather than "is the box alive", plus the on-call and post-incident routine that turns outages into fixed causes.

Assessment3 to 6 weeks

Reliability baseline, SLO proposal, load-test findings, prioritised plan.

AdvisoryFractional authority alongside the platform team

SLO set, error-budget policy, arbitration on headroom spend.

Interim seat6 to 18 months, CTO or COO-adjacent

The practice stood up and the service level held through change.

SLO and error budgetsLoad and chaos testingCapacity modellingObservability and SREPerformance budgets in CI

State of the proof

counted from the ledger below
4sub-capabilities
1proven · a published case carries a sourced figure
3held · carried by a named operator, no case published
0declared · in scope, no published proof today

Experiences that prove it

1 detailed mandate

Sub-capabilities and evidence

4 lines, each with its state
Reliability engineering and availability targetsprovenPMU
Application and platform performanceheldResponse-time and load work carried inside delivery mandates — a performance budget wired into the pipeline — with no case published on the performance work in its own right.
Capacity and scalability planningheldCapacity modelling held on high-availability estates where peak load is the design constraint; not published separately from the migrations that carried it.
Observability and incident responseheldSignal design, on-call rotation and post-incident practice run inside platform mandates; reported as part of the migration record rather than as a standalone engagement.

Publishable figures

named, sourced, attributable
99.995%Service-level uptime held through the migrationPMU mandate · five-nines target on live stakes
-87%TCO cut without giving up the service levelPMU mandate · replatformed betting applications

Questions

answered, in the open
Why hold performance next to cost?
Because headroom is bought twice — in the invoice and in the footprint. An SLO with no cost owner produces over-provisioning; a cost programme with no SLO produces outages. The trade-off is argued once, in the open, by someone accountable for both.
Where is the real proof?
In the PMU replatforming, where a five-nines 99.995% uptime target held on a live betting platform while total cost of ownership fell 87% — the service level was the constraint the cost work had to respect.
What is honestly not published?
Standalone performance engagements. Response-time, capacity and observability work is held by named operators and carried inside delivery and migration mandates; there is no published case sold as performance work on its own.
Is this the same as platform engineering?
No. Platform engineering builds the paved road; this werk is accountable for how fast it is and whether it stays up. They are frequently bought together and they answer to different numbers.

Related werks

same family, shared proof
Discuss a perfops mandateHow Advisory runs itAll sixteen werks