Reliability › Observability

Observability

What a NestLaravel service tells you about itself, how to switch it on, and what it deliberately does not do. Everything below is provided by nestlaravel/kafka and is on by default in generated services unless noted.

#1. Structured logs

Use the nestlaravel log channel (LOG_CHANNEL=nestlaravel, set in the generated Docker/Compose env). Each line is one JSON object on stdout:

{class="tk-s">"ts":class="tk-s">"2026-09-30T13:05:11+00:00",class="tk-s">"level":class="tk-s">"warning",class="tk-s">"message":class="tk-s">"Saga step failed",class="tk-s">"service":class="tk-s">"orders",
 class="tk-s">"request_id":class="tk-s">"9f1c…",class="tk-s">"correlation_id":class="tk-s">"ord_123",class="tk-s">"causation_id":class="tk-s">"e5d4…",class="tk-s">"event_id":class="tk-s">"e5d4…",
 class="tk-s">"tenant_id":class="tk-s">"acme",class="tk-s">"trace_id":class="tk-s">"4bf92f3577b34da6a3ce929d0e0e4736",class="tk-s">"span_id":class="tk-s">"00f067aa0ba902b7",class="tk-s">"context":{…}}

#2. Metrics (Prometheus)

GET /metrics returns Prometheus text format, protected by a bearer token (METRICS_TOKEN). If the token is unset the endpoint returns 404 (fail closed). Counters/histograms live in the cache store METRICS_CACHE_STORE (use Redis in production so all PHP-FPM workers and daemons aggregate; with file/array each process only sees itself).

AreaMetrics
HTTP (RED)nestlaravel_http_requests_total, nestlaravel_http_request_duration_seconds, nestlaravel_http_errors_total
Kafka consumernestlaravel_kafka_consumed_total{result}, nestlaravel_kafka_processing_seconds, nestlaravel_kafka_retries_total, nestlaravel_kafka_dlq_total, nestlaravel_kafka_consumer_errors_total, nestlaravel_kafka_commit_failures_total
Inboxnestlaravel_inbox_total{result=processed or duplicate}
Outboxnestlaravel_outbox_published_total, _retries_total, _failed_total, _recovered_total, _lost_claims_total; gauges nestlaravel_outbox_pending, _processing, _failed, _oldest_pending_age_seconds; nestlaravel_events_produced_total
Upstream callsnestlaravel_upstream_requests_total, nestlaravel_upstream_seconds, nestlaravel_upstream_retries_total, nestlaravel_circuit_rejected_total, nestlaravel_circuit_transitions_total
Saganestlaravel_saga_total{saga,result}
Platformnestlaravel_database_up, nestlaravel_db_query_duration_seconds (opt-in), nestlaravel_queue_depth, nestlaravel_jobs_total, nestlaravel_cache_errors_total, nestlaravel_process_memory_bytes, nestlaravel_process_memory_peak_bytes, nestlaravel_host_load

Where the numbers live matters. Metrics are kept in a Laravel cache store (METRICS_CACHE_STORE, default: the app's default store). Laravel's default store is database, where every metric write is a SQL query; file is a file write; array is lost at the end of each request. Use Redis (METRICS_CACHE_STORE=redis) in production; nestlaravel production:check warns otherwise. A histogram observation costs 3 cache writes, a counter 1. Two consequences of this design, both found and fixed after a clean-install test failed: metric writes are re-entrancy guarded (they can never measure themselves), and counters work on the database store (which, unlike Redis/array, does not create a key on increment). nestlaravel_db_query_duration_seconds is opt-in (METRICS_DB_QUERIES=true): one observation per SQL query is measurable overhead.

Metrics never contain per-request ids, user ids or tenant ids as labels (cardinality); those belong in logs and traces. Metrics are best-effort: a failing cache never fails a request or a message (it increments nestlaravel_cache_errors_total where it can).

#Alerts worth having

SymptomExpression (sketch)
Events are not leaving the databasenestlaravel_outbox_oldest_pending_age_seconds > 60
Publishing is failing permanentlyincrease(nestlaravel_outbox_failed_total[10m]) > 0
Consumers are parking messagesincrease(nestlaravel_kafka_dlq_total[10m]) > 0
Offsets not committingincrease(nestlaravel_kafka_commit_failures_total[10m]) > 0
A dependency is being shedincrease(nestlaravel_circuit_transitions_total{to="open"}[5m]) > 0
A saga needs a humanincrease(nestlaravel_saga_total{result="failed"}[5m]) > 0
Database unreachablenestlaravel_database_up == 0

#3. Tracing (optional)

#4. Health endpoints

EndpointMeaningDepends on
GET /livenessProcess is up and can answer. Restarting fixes nothing else.nothing
GET /startupBoot finished: database reachable, reliability tables migrated.database
GET /readinessSafe to receive traffic.only the dependencies in HEALTH_REQUIRED (default database)
GET /healthFull report per dependency; degraded when optional ones are down (HTTP 200), down (503) when a required one is.all

Kafka is not in the default readiness set on purpose: a broker outage must not remove every pod from the load balancer, because HTTP requests still succeed (events queue in the outbox). Add kafka to HEALTH_REQUIRED for pure consumer workers.

#5. Not included

Dashboards, alert routing, log shipping, long-term storage, and a tracing backend. The endpoints and formats above are standard so any Prometheus / Loki / Elastic / Tempo / Jaeger / Datadog stack can consume them.

Edit this page on GitHub