Know what production is doing
Observability
Caspian ships request correlation, structured request and event logs, fixed-cardinality process
metrics, and /health / /ready probes. Development
diagnostics in the terminal and the browser log stay separate.
Configuration
Settings come from the environment with production-on defaults. Logs go to stderr — container runtimes collect them without filesystem access, so no log file is written.
| Variable | Default | Effect |
|---|---|---|
| REQUEST_LOGS | on in production | One record per completed request |
| LOG_JSON | on in production | JSON lines on stderr instead of readable text |
| SLOW_REQUEST_MS | 1000 | Slow threshold; 0 disables the flag |
| LOG_HEALTH_CHECKS | off | Include probes in records and metrics |
| READINESS_TIMEOUT_SECONDS | 2 | Per-check timeout for /ready |
| SERVICE_NAME | project name | The service field |
| REQUEST_ID_TRUST_INCOMING | on | Keep a valid incoming X-Request-Id |
Request records
{"kind":"request","timestamp":"2026-09-15T14:02:11.482Z","level":"info","service":"shop", "method":"GET","path":"/orders/42","status":200,"durationMs":3.42,"slow":false,"requestId":"84d9…"}
- The path never includes the query string. Bodies, cookies, headers, CSRF tokens, and uploads are never recorded.
-
leveliserrorfor 5xx,warnfor 4xx or slow requests,infootherwise. -
durationMsmeasures time until the response started; streams and downloads stay open longer. -
Existing
public/files are not recorded; maintenance and rate-limit refusals are.
Application events
from casp.observability import record
record("info", "checkout", "order accepted", orderId=order.id)
{"kind":"event","timestamp":"…","level":"info","service":"shop","target":"checkout", "message":"order accepted","requestId":"84d9…","fields":{"orderId":42}}
Inside a request the current request id is added automatically. Keyword fields nest under
fields so they cannot overwrite the envelope; level, target, and message are truncated at 16 KiB.
Server errors handled by Caspian use this stream with target http and include the
traceback in fields.trace — server-side only; the client response stays redacted.
The application owns every field. Never record passwords, tokens, cookies, session payloads, or personal data the log destination is not approved to hold.
Liveness and readiness
GET /health
200 while the process serves HTTP. Checks nothing, on purpose: restarting a healthy process does not repair an unavailable database.
GET /ready
200 when every registered check passes, otherwise 503. Checks run concurrently and report ok, failed, or timeout.
{"status": "unavailable", "checks": [{"name": "database", "status": "timeout"}]}
Both probes are no-store and bypass route privacy, rate limiting, maintenance mode, and the
application pipeline. Do not create src/app/health or src/app/ready
routes. When the generated Prisma client exists, a database check (SELECT 1)
is registered automatically. Add your own in src/lib/readiness.py:
# src/lib/readiness.py
from casp.observability import readiness_check
@readiness_check("payments")
async def payments_ready():
await payments_client.ping()
A check may be sync or async and passes unless it raises, returns False, or times out. Names are 1–64
characters and unique. Keep checks cheap and side-effect free: verify a required dependency — never migrate, create records,
consume messages, or call an optional service. Failure details go to the log stream, never the response.
Metrics
from casp.observability import metrics metrics() # {"requests": 1840, "active": 3, "client_errors": 21, "server_errors": 0, # "slow": 4, "duration_total_us": 5120331, "duration_max_us": 812004}
Counters are process-local, reset on restart, and carry no route, user, or status labels, so attacker-controlled paths cannot grow them. There is no public metrics route — export through an approved collector or an authenticated internal endpoint, and aggregate worker snapshots in your monitoring system.
Operational boundary
Caspian provides ids, records, probes, and counters. Your deployment owns log collection, retention, redaction policy, alerting, dashboards, trace export, and fleet aggregation.
Error handling and request ids