You are asked "is this service production-ready?" What checklist do you walk through?
basicProduction-ready means it can be deployed safely, observed, operated by someone else at 3am and fails gracefully. Walk these areas:
- Build and deploy: reproducible image, CI gates, rollback tested.
- Config and secrets externalized; no secrets in the image or Git.
- Health probes, graceful shutdown, resource requests/limits, at least 2 replicas with a PodDisruptionBudget.
- Resilience: timeouts, retries with backoff, circuit breakers, idempotency.
- Data: backward-compatible migrations, backups, connection pool sizing.
- Observability: metrics, structured logs, traces, SLO alerts with runbooks.
- Security: authn/authz, TLS, dependency/image scanning, least-privilege.
- Ownership: on-call, dashboards, capacity estimate, load test.
- Is passing unit tests production-ready? No, nothing there covers failure modes, operations or load.
- Who signs off? Ideally an automated readiness review plus a human review for new services.