Reliability › Failure Scenarios

Failure scenarios

For every dependency: what the caller sees, what the system does, how it recovers, and which automated test proves the logic (tests use SQLite + an in-memory broker; see RELIABILITY.md for what that does and does not prove). "Drill" means: reproduce it in staging with the real component.

Legend: ✅ proven by an automated test · 🔧 provided by config/manifests, not exercised by tests · ⚠️ your responsibility

#1. Kafka broker down

client ─▶ gateway ─▶ orders ──tx──▶ [orders + outbox row]   ✅ HTTP 2xx, event durable
                                        │ publisher: produce fails → backoff 30s·2ⁿ (cap 15m) → status stays retryable
                                        ▼
                                  broker returns → backlog drains → consumers' inbox absorbs any duplicates

#2. Consumer crashes

receive ─▶ BEGIN ─ inbox INSERT ─ handler writes ─ COMMIT ─▶ [CRASH] ─▶ offset not committed
restart ─▶ redelivery ─▶ inbox INSERT ignored (duplicate) ─▶ handler NOT run ─▶ commit offset

#3. Outbox publisher crashes

#4. Database down or slow

request ─▶ DB error ─▶ 500/503 (no partial writes: event + data commit together)
/liveness 200 (no restart storm)   /readiness 503 (drained from LB)   consumer: retry ×3 → DLQ (parked, partition moves on)

#5. Redis down

FunctionBehaviour
Application caches (SafeCache)bypass to source of truth for a short window, error counter ✅ DegradedModeTest
Circuit breakerfails open (traffic flows) ✅
HMAC replay protectionfails closed: 503 to the caller, because accepting unverifiable requests would silently disable replay protection ✅ GatewaySignatureTest. Opt out with INTERNAL_REPLAY_PROTECTION_REQUIRED=false if availability matters more
Metricsdropped; requests unaffected
Rate limit / sessions / Redis queuesunavailable: use Laravel's failover drivers or accept the outage ⚠️

#6. Downstream service down (gateway → service)

gateway ─timeout 3s connect / total─▶ service
   transient (conn error, 408/425/429/5xx) & safe method (or Idempotency-Key) → retry with backoff, re-signed
   N failures → breaker OPEN (shared in cache, all replicas agree) → 503 + Retry-After without calling the service
   half-open probe → close on success

#7. Poison message

Malformed JSON, missing envelope fields, schema violation, unsupported event_version, or NonRetryable → straight to <topic>.dlq, handler never runs, offset committed, partition keeps moving ✅ PipelineReliabilityTest, SchemaGovernanceTest, ChaosScenariosTest::test_malformed_event_on_the_topic_does_not_block_the_partition. If even the DLQ publish fails the consumer exits without committing – the message is never silently dropped ✅. Inspect with nestlaravel dlq:list <topic>.

#8. Long-running workflow fails halfway

Saga compensation in reverse order, timeouts, retries and crash recovery ✅ SagaTest (see SAGA.md). A compensation that keeps failing parks the saga as failed and logs at critical.

#9. Tenant context lost

Queued jobs carry tenant_id; events carry it in the envelope; logs always include it. With TENANCY_STRICT_JOBS=true dispatching a job with no tenant is refused ✅ TenantPropagationTest. nestlaravel tenant:check audits models vs tables ✅ TenantCheckTest. ⚠️ Raw DB::table() queries bypass model scopes; review them.

#10. Deploy / restart

Rolling updates with maxUnavailable: 0, /startup gating, preStop delay and terminationGracePeriodSeconds 🔧 (infrastructure/k8s, validated by kubeconform). Consumers and the outbox daemon stop cleanly on SIGTERM ✅ (Linux).

#What is not covered by any test in this repository

Network partitions between services and the broker, broker rebalance storms, disk-full on the database, clock skew beyond the HMAC timestamp window, and multi-region failover. Rehearse them with the drills in OPERATIONS.md.

Edit this page on GitHub