Skip to main content
Start a Conversation

What to Monitor Before Production Launch

Start with the customer task, then trace the dependencies that make it work.

Shawn Iuliucci
4 min read
Reliability & Platform Engineering
On this page

A production dashboard that shows CPU and memory but cannot detect a failed inquiry or order is incomplete. Identify the user action the system exists to support, measure success and failure, and add dependency signals that help an operator diagnose the result.

Choose user-facing indicators

Track completed transactions, error rates, latency, and queue age for important workflows. Define what an actual customer would notice before setting alert thresholds.

Connect traces and logs

Carry a safe correlation identifier across services so a failed request can be followed without exposing sensitive payloads in logs.

Make alerts actionable

Route each alert to an owner and a runbook. Alert on symptoms that need intervention, not every transient fluctuation in an internal metric.

Define service signals from user journeys

Choose the few actions customers or staff need most: submitting an inquiry, logging in, placing a job, or receiving a result. For each, measure completion, failure, delay, and dependency state. A synthetic check can test the path, but it should use a controlled account and avoid creating noisy production records. Connect transaction identifiers across services so an operator can trace a failure without logging private payloads. Infrastructure metrics still matter, but they become diagnostic clues tied to a customer-visible symptom.

Make a journey sheet for the top customer actions. Each row should show the entry point, success event, failure event, expected duration, downstream dependency, and alert owner. A form journey, for example, may include stored submission and confirmed CRM delivery as separate milestones. Put a safe synthetic check beside any path that can fail silently. Link each alert to a runbook and define the symptom that merits waking a person. When a release changes the flow, update the sheet and test the signal before declaring the release done. This ties observability to service behavior instead of dashboard decoration.

Practice what happens when an alert fires

Give each actionable alert an owner, threshold rationale, first check, and escalation route. Test an intentionally failed dependency to confirm the alert reaches the right person and shows enough evidence to act. Review false alarms and missing alerts after launch. Keep dashboards for current behavior, not only historical capacity. If a form is accepted locally but delivery to the CRM stalls, the system should show both facts and the age of the queue. A monitoring setup is ready when the team can detect, explain, and recover from a representative failure.

Decision checklist

  • List the top three customer journeys.
  • Create a synthetic or controlled check for one critical path.
  • Test alert delivery and ownership.
  • Review logs for sensitive data before launch.

A small test before committing

Pick three customer tasks, such as submitting a form, completing a transaction, and loading a dashboard. Run them in a test environment while deliberately breaking a dependency, slowing a database call, and returning an invalid downstream response. Confirm that the alert identifies the affected task and reaches an owner before a customer report is needed. If monitoring only says the server is up, add task-level signals and a response path.

Worked scenario

A hypothetical website may return HTTP 200 while its form-to-CRM handoff is failing. Monitoring should distinguish accepted website submissions, pending deliveries, and confirmed downstream receipt so the team knows what to repair.

For a scoped application of this decision, see Reliability & Platform Engineering.

Apply this decision to your own system.

Share your current workflow and constraints so the next step can be scoped around real work.

Discuss Your Project