Reliability › Guarantees & Tests

Reliability

What NestLaravel guarantees between services, how, and — just as important — what it does not guarantee. Every guarantee below names the automated test that demonstrates it (packages/laravel-kafka/tests/Reliability/*, apps/*-service/tests, apps/api/tests). If there is no test, it is listed under Remaining risks, not claimed.

#Delivery semantics (the honest summary)

 Business transaction (service A)                     Consumer (service B)
 ┌──────────────────────────────┐                     ┌──────────────────────────────────────────┐
 │ UPDATE orders …              │                     │ receive → validate envelope + schema     │
 │ INSERT outbox_messages(event)│  atomic (same tx)   │ BEGIN                                    │
 └───────────────┬──────────────┘                     │   INSERT inbox(consumer,event_id) -- unique
                 │ COMMIT                             │   0 rows? → duplicate → COMMIT → ACK     │
   outbox publisher (claim → produce → flush)         │   handler()   -- your writes             │
                 │ broker-confirmed                   │ COMMIT                                   │
                 ▼                                    │ commit Kafka offset  (only now)          │
              Kafka  ─────────────────────────────▶   └──────────────────────────────────────────┘

#Idempotent consumption — the inbox

class="tk-c">// inside a MessageHandler (the consumer pipeline does this for you when KAFKA_INBOX_ENABLED=true)
app(EventInbox::class)->process(class="tk-v">$event[class="tk-s">'event_id'], fn () => class="tk-v">$this->reserveStock(class="tk-v">$event));
FailureResultTest
Same event delivered twicesecond delivery is skipped, one business effectInboxTest::test_duplicate_delivery_produces_one_business_effect
Handler throwsbusiness writes and dedup record roll back; retry runs the handler once…test_handler_exception_rolls_back_business_and_dedup_record_so_retry_succeeds
DB error inside handlereverything rolled back…test_db_failure_inside_handler_rolls_everything_back
Second worker gets the event while the first is mid-transactionsecond waits on the unique index, then skips (re-entrant variant tested; multi-process variant runs on PostgreSQL in CI)…test_concurrent_duplicate_inside_a_running_handler_is_skipped, ConcurrentInboxTest
Crash after DB commit, before Kafka offset commitevent redelivered, deduplicated, no duplicate payment…test_consumer_crash_after_commit_before_offset_commit_is_harmless, PipelineReliabilityTest::test_redelivery_after_crash_between_db_commit_and_offset_commit_charges_once
Retry after a transient failuresucceeds exactly once…test_transient_handler_failure_is_retried_and_succeeds_once

The inbox table (inbox_events, unique (consumer, event_id)) must live in the same database as the business tables. Retention (KAFKA_INBOX_RETENTION_DAYS, default 14; php artisan inbox:prune) must exceed the longest possible redelivery window (topic retention, consumer downtime).

#Transactional outbox

Events are written to outbox_messages inside the caller's transaction (OutboxEventBus) — the business change and the event commit or roll back together (OutboxReliabilityTest::test_business_row_and_outbox_row_commit_atomically).

pending ──claim──▶ processing ──flush ok──▶ published
   ▲                   │
   │ backoff           ├─ failure, attempts < max ─▶ pending (available_at = now + base·2^(n-1), capped)
   └───────────────────┤
                       ├─ failure, attempts ≥ max ─▶ failed  (+ copy on <topic>.dlq, last_error kept)
                       └─ publisher died (locked_at older than visibility timeout) ─▶ pending

#Consumer failure handling

kafka:consume (test: ConsumerFailureTest, PipelineReliabilityTest)

SituationBehaviour
Handler throwsretried KAFKA_CONSUMER_MAX_RETRIES times with exponential backoff, then DLQ (<topic>.dlq)
Handler throws NonRetryablestraight to the DLQ
Malformed JSON / bad envelope / schema violation / unsupported version (poison)straight to the DLQ, handler never runs
Offset commitonly after processing or dead-lettering; never before
DLQ publish failsconsumer exits without committing — the message is neither lost nor skipped
Offset commit failslogged + counted, message will be redelivered and deduplicated
Broker unavailable / network flaptransient: exponential backoff, keeps running (test_transient_broker_errors_back_off…)
Authentication / TLS / ACL errorfatal: exits 1, does not hammer the broker
Topic missingclassified transient; set KAFKA_CONSUMER_MAX_CONSECUTIVE_ERRORS to escalate
SIGTERMfinishes the in-flight message, commits, leaves the group, closes DB/Redis, exits 0
Rebalancehandled by librdkafka; unacknowledged messages are redelivered (inbox absorbs duplicates)

Knobs: retry count/backoff, --timeout, --max, --max-runtime, --memory. Concurrency = number of consumer processes in the same group (≤ partitions). Batch size at the consumer is 1 by design (offset safety).

#Event contract & schemas

{ class="tk-s">"event_id": class="tk-s">"uuid", class="tk-s">"event_type": class="tk-s">"orders.order.created", class="tk-s">"event_version": 1, class="tk-s">"occurred_at": class="tk-s">"…",
  class="tk-s">"producer": class="tk-s">"orders-service", class="tk-s">"correlation_id": class="tk-s">"uuid", class="tk-s">"causation_id": class="tk-s">"uuid|null", class="tk-s">"tenant_id": class="tk-s">"…|absent",
  class="tk-s">"traceparent": class="tk-s">"00-…-…-01|absent", class="tk-s">"payload": { } }

#Sagas (long-running workflows)

See SAGA.md. Persisted state, idempotent start/resume, reverse-order compensation, timeouts, crash recovery (SagaTest, 12 tests).

#Timeouts, retries and circuit breaking (gateway → service)

ResilientHttp (ResilienceTest, GatewayResilienceTest)

#Failure modes of shared infrastructure

Dependency downWhat happensTest
Kafka brokerHTTP keeps working; events accumulate in the outbox (durable), published when the broker returns; readiness stays green unless you list kafka in HEALTH_REQUIRED; consumers back offtest_broker_down_keeps_rows…, test_readiness_fails_only_for_required_dependencies
Redisoptional caches degrade to the source of truth (SafeCache); circuit breaker fails open; replay protection fails closed (503) unless you opt out; rate limits/queues/sessions on Redis are unavailable — configure Laravel's failover cache/queue driver or accept the outageDegradedModeTest, test_requests_are_refused_when_replay_protection_cannot_be_guaranteed
Databasereadiness → 503 (traffic drained), liveness stays 200 (no restart storm); consumers retry then DLQ; outbox rows stay durabletest_readiness_reports_a_down_database
Gateway replicastateless; breaker/rate-limit/nonce state in shared cache; any replica can serve any requesttest_breaker_state_is_shared_through_the_cache_not_process_memory
Payment service crashes after DB commit, before offset commitredelivery → inbox → no duplicate paymentPipelineReliabilityTest, ChaosScenariosTest

#Database

#Remaining risks

Not claimed, because not proven by automated tests in this repository:

  1. Real multi-process inbox concurrency is tested in CI against PostgreSQL 16 and MySQL 8.4 (ConcurrentInboxTest, separate OS processes, CI job "Reliability on real databases"; it fails the build if skipped). MariaDB, other isolation levels, replicas / failover and cluster setups are not tested; SQLite (unit tests) serialises writers and proves only the logic.
  2. End-to-end with a real broker covers produce → consume → DLQ (RdKafkaBrokerTest) in CI; broker restarts, network partitions and rebalance storms are not automated (documented drills in OPERATIONS.md).
  3. Graceful shutdown is tested for the consumer and outbox daemon on Linux CI (GracefulShutdownTest: a real SIGTERM while a message/batch is in flight). Those tests are skipped without ext-pcntl (Windows), so they do not run on a Windows dev machine. HTTP drain depends on your proxy/orchestrator settings (provided in infrastructure/k8s, validated as manifests by kubeconform, never deployed to a cluster by CI).
  4. OpenTelemetry export is an OTLP/HTTP JSON exporter verified against a fake collector, not against a real Collector/Jaeger/Tempo.
  5. Kafka ACLs, TLS certificate rotation, broker replication are operator responsibilities and are not verifiable from inside the application (production:check says so).
  6. Backups / restore procedures are documented but the framework cannot test your restore (DISASTER-RECOVERY.md).
  7. Outbox publisher should run as a single instance per service for strict per-aggregate ordering; multiple publishers are safe (no double publish) but may reorder events of one aggregate.
  8. Performance overhead is measured on a developer machine with SQLite (BENCHMARKS.md); production numbers depend on your database and broker.
  9. The reference application and the chaos scenarios (REFERENCE-APP.md, ChaosScenariosTest) run on an in-memory broker and one shared SQLite database. They prove the logic of crash/duplicate/compensation handling end to end, not the behaviour of real independent databases, a real cluster, or network faults.
  10. Effects outside your database (payment providers, e-mail) are outside every transaction here; only idempotency keys on those calls make them safe to repeat.
Edit this page on GitHub