Data and analytics + Platform and reliability

Data platforms and pipelines

Your numbers live in five tools, a few spreadsheets and one script on someone's laptop. Reports break when a column changes, and nobody notices until a meeting. We build the platform that collects, models and checks your data, and we run it with the same discipline as the systems that make you money.

What we do

Get the data in one place. We pick the sources that matter, such as your product database, billing, CRM and ad platforms, and load them into a warehouse or a Postgres analytics database. Raw data lands first and stays available. We use managed connectors where they are cheaper than writing code, and custom loaders where they are not.

Model it once, in SQL. Transformations live as SQL models in version control, in a dbt-style layout: staging, then clean business tables, then the few tables your dashboards read. Every model has tests and a short description. When a definition changes, it changes in one file and everyone gets the new number.

Run it like production. Pipelines get an orchestrator, retries, and a freshness target per table. Data quality checks catch duplicates, nulls and sudden drops in row counts. When something breaks, an alert reaches someone who can fix it, with a runbook attached. This is the same SRE practice we apply to customer-facing systems, because a wrong number in a board meeting is also an outage.

Keep the bill sane. Warehouses are easy to overspend on. We look at which queries and jobs cost the most, set partitions and schedules that match how the data is used, and give you a cost view per source and per job.

Fit the platform to the company. A ten-person team does not need the stack of a bank. We start with the smallest setup that answers your questions, often a managed Postgres and a scheduler, and move to a cloud warehouse only when volume or the number of users makes it worth it. Everything is defined as code, so the next step is a change, not a rebuild.

How it usually goes

Most engagements start with the two-week data health check. We trace your key numbers from source to dashboard and find where they drift or go missing. That gives us a clear list of what the platform has to fix first.

Then we build in two-week iterations, a few sources at a time. The same team sets up the infrastructure and writes the analysis on top, so there is no hand-off between a platform vendor and a data vendor. When we leave, your team owns the code, the configs and the runbooks.

A good fit if

  • Your reports pull straight from production databases and slow them down.
  • Two teams quote two different numbers for the same thing.
  • A pipeline broke last month and nobody knew for a week.
  • You are about to hire your first analyst and want a clean place for them to work.

Questions we get

Do we need a data warehouse?
Not always. If your data fits comfortably in a single Postgres instance and comes from a few sources, a separate analytics schema or a read replica is often enough. A cloud warehouse such as BigQuery or Snowflake pays off when data volume, the number of sources or concurrent analysts grow past what one database handles well. We start with the smallest setup that works and leave room to move.
What is the difference between ETL and ELT?
ETL transforms data before loading it into the warehouse. ELT loads the raw data first and transforms it inside the warehouse with SQL. ELT is the usual choice today because raw data stays available, transformations are versioned and testable, and you can rebuild a model without going back to the source.
What is dbt and do we need it?
dbt is a tool for writing data transformations as SQL models, with tests, documentation and dependencies between models. It is a good default for teams that already know SQL. If you only have a handful of transformations, plain SQL views with a scheduler can be enough, and we will tell you so.
How do you keep data pipelines reliable?
We treat them like production services. Each pipeline has an owner, a freshness target, automated data quality tests and alerts that go to a person, not a channel nobody reads. Failures get a short review and a fix, the same way a production incident would.

Tell us what's broken.

A few sentences is enough. We reply within one working day and the first call is free.