| Transient API failure | An instance stops mid-process because a connected system briefly returned an error or a timeout. | Bounded retry with a delay, distinguishing retryable responses from permanent ones, and a defined give-up path that notifies rather than sits. |
|---|
| Throttling under volume | Fine in testing, failing in month-end. Calls are being rejected because too many were made too quickly. | Move calls out of loops where possible, batch where the API supports it, pause between iterations, and respect the documented limits rather than discovering them. |
|---|
| Approver unavailable | A task sits for two weeks because the assignee left, is on leave, or is also the requester. | Due dates, reminders, an escalation path with a defined timeout, delegation rules, and a routing rule that never assigns a task to its own requester. |
|---|
| Partial multi-system write | The CRM was updated, the ERP was not. Two systems now disagree and nobody knows which is right. | Order writes so the authoritative system goes last where possible, and define a compensating action — plus a reconciliation check that flags the disagreement. |
|---|
| Self-triggering loop | A workflow updates a record, which re-triggers the same workflow, which updates the record. | A trigger condition that excludes the workflow's own service identity or its own field changes, plus an instance-count alert that catches it in test. |
|---|
| History flooding | Instances slow markedly and the history is unusable, because logging was placed inside a loop that now runs thousands of times. | Log outside loops, log decisions rather than iterations, and keep the history readable enough to be used during an incident. |
|---|
| Silent stall | Nobody notices an instance has stopped until a customer asks. There is no alert because the error was captured and discarded. | Every captured error routes to a named owner, and a stalled-instance check runs on a schedule so the process monitors itself. |
|---|