VISM does not replace operational judgment during a failure. Its resilience model is designed to fail closed for new mutation, preserve uncertainty, reconcile late effects and make recovery authority explicit.
Representative failure modes
| Failure | Why it matters | Designed response |
|---|---|---|
| Invented capacity claim | An agent can over-scale from unsupported reasoning. | Require deterministic capacity evidence; simulate or stop when missing. |
| Shared dependency omitted | Two services can consume the same budget. | Provenance graph, reservations and staged rollout. |
| Backup succeeded but restore fails | Backup completion is not recovery proof. | Checkpoint 1 cannot be verified without restore evidence. |
| Timeout after provider write | Blind retry can duplicate effect. | Persist intent, reconcile state/receipt and verify before next action. |
| Approval payload substitution | Approved prose can differ from sent effect. | Plan/permit/payload hashes and admission validation. |
| Old worker resumes | Lease expiry alone cannot stop a stale credential. | Fencing epoch at the effect sink and narrow credentials. |
Three integration exercises
- Scale exceeds dependency: a 3 → 1,000 request with a database ceiling of 240 should not dispatch an unsafe plan; a later stage failure should stop and preserve recovery evidence.
- Approval substitution: approve image digest A, then change payload, threshold or adapter before dispatch; every control boundary should deny the mismatch.
- Crash after effect: allow provider apply, lose the response and restart the coordinator; the workflow should reconcile without a duplicate scale.
Residual risk
Known residuals include unknown dependencies without telemetry, privileged customer administrators changing controls, compromised upstream IdP/KMS, correlated model errors, false application health, uncancellable provider side effects and delayed business damage. Mitigation is constrained scope, independent probes, staged execution and human incident response—not an absolute safety claim.
Questions that can narrow or stop the product
The product should not broaden when it cannot close direct-write bypasses, fence stale workers at the effect sink, bind approval to exact effect, keep verification independent, quantify data-loss recovery, or establish buyer willingness to pay. These are product and safety gates, not a backlog to conceal with a more capable model.
GO / NARROW / STOP
Proceed when enforcement works, a scoped pilot has measurable value, owners trust recovery and paid demand appears. Narrow to verification/evidence when only one use case is viable. Stop when a hard safety assumption or willingness-to-pay test fails in the target customer profile.