Engineering blog

Scaling Postgres to 40k writes per second

Lessons from a Shopify-scale checkout rewrite.

GitOps for regulated workloads

Auditable deployments without locking developers out.

Why we run the same stack in 4 regions

A practical guide to multi-region Kubernetes.

On-call without burnout

Error budgets, follow-the-sun and blameless reviews.

Serving LLMs at scale with vLLM

GPU scheduling, autoscaling and cost control in production.

The observability stack we install for everyone

Prometheus, Grafana, Loki and OpenTelemetry done right.

Migrating 200 services without downtime

A playbook for preserving customers' trust during a move.

Internal developer platforms: worth the build?

We built one three times. Here is when it pays off.