DOCUMENTATION

Resilience

Failure-mode thinking, integration exercises and conditions that require narrowing or stopping.

Docs menu

VISM does not replace operational judgment during a failure. Its resilience model is designed to fail closed for new mutation, preserve uncertainty, reconcile late effects and make recovery authority explicit.

Representative failure modes

FailureWhy it mattersDesigned response
Invented capacity claimAn agent can over-scale from unsupported reasoning.Require deterministic capacity evidence; simulate or stop when missing.
Shared dependency omittedTwo services can consume the same budget.Provenance graph, reservations and staged rollout.
Backup succeeded but restore failsBackup completion is not recovery proof.Checkpoint 1 cannot be verified without restore evidence.
Timeout after provider writeBlind retry can duplicate effect.Persist intent, reconcile state/receipt and verify before next action.
Approval payload substitutionApproved prose can differ from sent effect.Plan/permit/payload hashes and admission validation.
Old worker resumesLease expiry alone cannot stop a stale credential.Fencing epoch at the effect sink and narrow credentials.

Three integration exercises

  1. Scale exceeds dependency: a 3 → 1,000 request with a database ceiling of 240 should not dispatch an unsafe plan; a later stage failure should stop and preserve recovery evidence.
  2. Approval substitution: approve image digest A, then change payload, threshold or adapter before dispatch; every control boundary should deny the mismatch.
  3. Crash after effect: allow provider apply, lose the response and restart the coordinator; the workflow should reconcile without a duplicate scale.

Residual risk

Known residuals include unknown dependencies without telemetry, privileged customer administrators changing controls, compromised upstream IdP/KMS, correlated model errors, false application health, uncancellable provider side effects and delayed business damage. Mitigation is constrained scope, independent probes, staged execution and human incident response—not an absolute safety claim.

Questions that can narrow or stop the product

The product should not broaden when it cannot close direct-write bypasses, fence stale workers at the effect sink, bind approval to exact effect, keep verification independent, quantify data-loss recovery, or establish buyer willingness to pay. These are product and safety gates, not a backlog to conceal with a more capable model.

GO / NARROW / STOP

Proceed when enforcement works, a scoped pilot has measurable value, owners trust recovery and paid demand appears. Narrow to verification/evidence when only one use case is viable. Stop when a hard safety assumption or willingness-to-pay test fails in the target customer profile.