Platform and reliability

Kubernetes and infrastructure as code

Kubernetes and Terraform are tools. They earn their complexity only when you need what they do. We set them up, clean them up, or tell you that a managed service would serve you better, and why.

What we do

Decide if Kubernetes is the right tool. We look at your workloads, your team and your budget. If a managed service or a few VMs would do the job, we say so, and help you get there. If Kubernetes fits, we set it up properly. That means managed control plane, sane node pools, network policies, workload identity, resource limits and a clear upgrade path.

Make upgrades boring. Many clusters run old versions because nobody wants to touch them. We write an upgrade strategy, test it in a staging cluster, clear out deprecated APIs and get you onto a regular rhythm. After that an upgrade is a normal task. It has a checklist, a rollback plan and an owner, and it does not need a weekend.

Write Terraform people can maintain. We design and review modules that stay readable rather than clever. Inputs that make sense, no hidden magic, and tests that run in CI. We split state so plans stay fast and a mistake in one area can’t break another. We also review existing code and tell you what to keep. Not every old module needs a rewrite, and a big-bang refactor of working Terraform is a risk of its own.

Close the gap between code and reality. Drift happens when someone fixes production by hand and never updates the repo. We find it, bring it back into code, and set up detection so it shows up early. Where continuous reconciliation pays off, we add GitOps with Argo CD or Flux, so the cluster always matches git. If a nightly plan and a pull request workflow are enough, we stop there.

How it usually goes

We usually start with a short review of your clusters and infrastructure code. We run plans, read modules, check versions and look for drift. Within two weeks you get a written report with what to fix, in order of risk.

Then we either guide your team through the fixes or do the first round with them, usually over four to eight weeks. We leave once your engineers can run an upgrade or change a module without us.

A good fit if

  • Your terraform plan takes so long that people skip it.
  • Your clusters no longer match what is in the repo.
  • Kubernetes upgrades are scary, so you keep putting them off.
  • You are about to adopt Kubernetes and want a second opinion first.

Questions we get

Do we need Kubernetes?
Many teams don’t. If you run a few services with steady traffic, a managed platform such as Cloud Run, App Engine, ECS or plain VMs is cheaper to run and easier to hire for. Kubernetes starts to pay off when you have many services, several teams and a real need for its scheduling, isolation and ecosystem.
GKE, EKS or self-managed Kubernetes?
Use a managed control plane unless you have a strong reason not to. GKE and EKS take care of the control plane and much of the upgrade work, which is where self-managed clusters cost the most time. On GKE, Autopilot removes node management too, at the price of some flexibility.
How should we structure Terraform state?
Split state by environment and by how often things change, so a plan only touches what it needs to. Keep it in a remote backend with locking, such as a GCS or S3 bucket. Avoid one giant state file and avoid splitting so fine that every change spans ten of them.
Argo CD or Flux?
Both do GitOps well. Argo CD has a strong web UI and suits teams that want to see sync status at a glance. Flux is lighter and fits well when everything is driven from git and the CLI. The bigger question is whether you need continuous reconciliation at all, and we answer that first.

Tell us what's broken.

A few sentences is enough. We reply within one working day and the first call is free.