Reliability › Operations Runbook

Operations runbook

How to run NestLaravel services in production and how to check that they are healthy. Pair with RELIABILITY.md (guarantees), OBSERVABILITY.md (signals) and DISASTER-RECOVERY.md.

#1. Process model

ProcessCommandReplicasNotes
HTTP (per service + gateway)php-fpm + nginx≥ 2, statelessgateway state (rate limit, circuit breaker, nonces) lives in the shared cache, so CACHE_STORE must be redis when replicas > 1
Outbox publisherphp artisan messaging:outbox-publish --daemon1 per service (strict ordering) — more are safe but may reorder one aggregate's eventsone deployment per service
Kafka consumerphp artisan kafka:consume <topic> <Handler> --max-runtime=3600 --memory=200≤ partitions of the topic, same KAFKA_GROUP_IDscale on consumer lag, not CPU
Saga recoveryphp artisan saga:recover every minute1 (scheduler)only for services that define sagas
Inbox retentionphp artisan inbox:prune daily1retention must exceed your longest redelivery window

Reference manifests: infrastructure/k8s/{service,workers,gateway}.yaml (validated with kubeconform in CI; not deployed to a cluster by CI). Compose: docker-compose.yml (stop_grace_period set for workers).

#Settings production should have (nestlaravel production:check reports on these)

APP_ENV=production, APP_DEBUG=false, KAFKA_ENABLED=true, KAFKA_SECURITY_PROTOCOL=ssl|sasl_ssl (FAIL if plaintext in production), producer acks=all + idempotence (FAIL otherwise), KAFKA_USE_OUTBOX=true, KAFKA_INBOX_ENABLED=true (the library default is false for upgrade compatibility – see UPGRADING.md; the generated Compose stack turns it on), KAFKA_SCHEMA_ENFORCE_PRODUCER=true and KAFKA_SCHEMA_ENFORCE_CONSUMER=true once every event has a schema, CACHE_STORE=redis (shared), DB_STATEMENT_TIMEOUT_MS > 0, METRICS_TOKEN set, INTERNAL_SERVICE_SECRET ≥ 32 chars, LOG_CHANNEL=nestlaravel. Things it can only remind you about (broker ACLs, backups, TLS termination) are reported as WARN "not verifiable".

#2. Health probes

ProbePathFailing meansOrchestrator action
startup/startupdatabase unreachable or reliability tables not migratedkeep waiting (do not send traffic)
liveness/livenessthe PHP process cannot answerrestart. Never depends on Kafka/DB/Redis, so an outage cannot cause a restart storm
readiness/readinessa dependency in HEALTH_REQUIRED (default database) is downremove from load balancer
full/healthJSON per dependency; degraded (200) vs down (503)dashboards / humans

The gateway keeps its own /health/live and /health/ready.

#3. Graceful shutdown

#4. Deploying

  1. Migrations first, additive only (expand → deploy → contract). Never rename/drop a column in the release that stops using it.
  2. Roll consumers before producers when an event schema gains a version; raise KAFKA_EVENT_MAX_VERSION only after all consumers understand it (php artisan events:check gates compatibility).
  3. Rolling update with maxUnavailable: 0. Wait for /readiness before shifting traffic.
  4. Verify: nestlaravel production:check, nestlaravel outbox:status, nestlaravel kafka:health.

#Rotating the service-to-service secret without downtime

  1. Set INTERNAL_SERVICE_SECRET_PREVIOUS=<old> and INTERNAL_SERVICE_SECRET=<new> on every service (services accept both, verifying the new one first). Roll them out.
  2. Switch the gateway to sign with <new>. Roll out.
  3. After all gateway replicas run the new value, remove INTERNAL_SERVICE_SECRET_PREVIOUS from services.

(GatewaySignatureTest: rotation, replay and expiry cases.)

#5. Day-2 commands

QuestionCommand
Are events leaving the DB?nestlaravel outbox:status (pending, processing, failed, oldest pending age)
Rows stuck in failed?nestlaravel outbox:status --failed, fix cause, then --requeue
Poison messages parked?nestlaravel dlq:list <topic>
Is the broker reachable?nestlaravel kafka:health
Which events/versions exist?nestlaravel events:list, nestlaravel events:check
Any cross-tenant exposure?nestlaravel tenant:check
Is this configuration sane?nestlaravel production:check (exit 1 on FAIL)

Replaying a dead letter: fix the root cause, then re-publish the original payload to the source topic; the inbox makes this safe even if part of it had been processed.

#6. Scaling

#7. Failure drills (run them in staging; the framework cannot run them for you)

Each drill: cause it, observe the listed signal, confirm the expected outcome, recover.

DrillExpected
Stop the Kafka broker for 5 min while creating ordersHTTP 2xx continues; outbox_pending and oldest_pending_age grow; after restart backlog drains, no missing or duplicate business effects
kill -9 a consumer during loadrebalance, redelivery, inbox duplicate counter increments, no double effects
Stop Rediscaches degrade; replay protection answers 503 (fail closed); breaker fails open; alert on nestlaravel_cache_errors_total
Stop the database/readiness 503, /liveness 200; recovery without restarts
Publish a malformed messagelands on <topic>.dlq, partition keeps moving
Deploy a schema-incompatible eventevents:check fails the pipeline

Automated equivalents of the in-process versions of these exist in ChaosScenariosTest; they use an in-memory broker and SQLite, so they prove the logic, not your infrastructure.

Edit this page on GitHub