Platform and reliability

Production readiness and SRE

You have a service heading to production, or already there, and someone has to answer the hard questions. What is the error budget? What wakes someone up at 3 a.m.? What happens when a zone goes away? We help you answer them before a customer does.

What we do

Define what “working” means. We pick the few user journeys that matter, such as signing up, paying, or loading the dashboard. Then we set service level objectives on them: how fast, how often successful, measured where the user feels it. These become the numbers your team argues with, instead of opinions.

Make alerts mean something. We cut alerts down to the ones that signal user impact or an error budget burning too fast. Everything else becomes a ticket or a dashboard. Fewer pages, and each one is worth getting up for.

Prepare for the bad day. We go through failure modes with your team: a zone outage, a full disk, a bad deploy, an expired certificate, a dependency that slows down instead of failing. Each gets a detection path, a runbook and, where possible, an automatic recovery. We test the risky ones on purpose.

Run incidents well, and learn from them. Clear roles during an incident, a status update rhythm, and a short blameless review afterwards that ends with owners and dates. Not a document nobody reads.

How it usually goes

Most engagements start with the two-week production readiness audit. We read your architecture, dashboards, alert history and the last few incident write-ups, and we spend time with the people on call. The audit ends with a fix list ordered by risk.

From there, some teams do the fixes themselves with a few hours of advice a month. Others ask us to do the first round with them, usually four to eight weeks, and hand over once the team is comfortable running it.

A good fit if

  • You are about to launch, or about to sign a customer with an SLA.
  • You just had an incident you don’t want to repeat.
  • Your team gets paged so often that they’ve stopped reading the alerts.
  • You have grown past the point where one person knows how everything works.

Questions we get

What is the difference between SRE and DevOps?
DevOps is a way of working: developers and operations share responsibility for running software. SRE is one concrete way to do it, with measurable targets (SLOs), an error budget that decides when to slow down, and engineering work to remove manual toil. We use SRE practices because they give you numbers to argue with instead of opinions.
Do we need a dedicated SRE team?
Usually not below 50 to 80 engineers. Most teams need clear SLOs, sane alerting and a shared on-call, owned by the product teams themselves. We set that up and train the people who will own it.
Which monitoring tools do you work with?
Prometheus, Grafana, Google Cloud Monitoring, CloudWatch, Datadog, OpenTelemetry and most of the usual suspects. We’d rather improve what you already run than migrate you to something new.
How long does a production readiness audit take?
Two weeks for a typical product with a handful of services. You get a written report with a fix list ordered by risk, and you keep it whether or not we do the fixes.

Tell us what's broken.

A few sentences is enough. We reply within one working day and the first call is free.