ENGINEERING BLOG · published Mon & Thu · 142 posts

How we run
one platform across 18 regions,
written by the people doing it.

Deep dives, post-mortems, design docs we wish we'd written before we shipped, and the occasional unhinged opinion. No PR-reviewed thought leadership. If we got it wrong, we say so.

All · 142 Engineering · 86 Post-mortems · 14 Design · 22 Open source · 11 Research · 9
Latest · 14 May 2026

02 ·Recent

Post-mortem·11 May 2026

The 14-minute eu-central-1 outage was actually three outages.

What our metrics said was one cascading failure. What we found after going through five days of trace data, log streams, and one extremely confused engineer.

Sara Khan · VP Reliability11 min
Engineering·06 May 2026

Predictive autoscale: how we forecast traffic 90 seconds out.

A 200-line model in Rust beats the reactive autoscaler we shipped two years ago on every metric we care about — including cold starts during the EU rush hour.

Dario Volpe · Engineer, ML Infra9 min
Design·28 Apr 2026

Why we ship a deploy as a content-addressed graph, not an artifact.

The design doc behind v4. How we got rollbacks to 4 seconds, regional fan-out to "free," and convinced our security team that immutable images are not a stretch goal.

Lia Marchetti · Principal Engineer14 min
$ scalable bench ▸ p50 ━━━━ 4.1ms ▸ p95 ━━━━━━━━━ 11.2ms ▸ p99 ━━━━━━━━━━━━ 18.7ms ✓ all good
Research·21 Apr 2026

Benchmarking 14 Postgres connection poolers, so you don't have to.

PgBouncer, PgCat, Supavisor, and 11 others — same workload, same hardware, 36 hours. The results were not what we expected, and one was disqualified for cheating.

Mira Chen · Data Platform22 min
Engineering·15 Apr 2026

Eight months of running our control plane on NATS: the good, the bad, the surprising.

We migrated from Kafka. Three things we miss, four things we don't, one thing we still can't believe works as well as it does.

Tomás Reyes · Distributed Systems16 min
−71% EGRESS COST · YoY
Engineering·08 Apr 2026

How we dropped customer egress bills by 71% without changing a single byte.

Just routing. Just BGP. Just patience. A diary of the 6 weeks we spent at three IXPs in Frankfurt, Marseille, and Singapore — and what came out of it.

Yuki Tanaka · Network Engineer13 min

03 ·Long-form

02 Apr2026
What it actually costs to run a 4-nines control plane in 18 regions.

The bill, the people, the on-call rotation. Not the marketing version — the engineering version, with line items.

EngineeringReliabilityEconomics
Sara KhanVP Reliability
·
28 min read
24 Mar2026
The single biggest mistake we made in v3, and how it shaped everything in v4.

A retrospective on the 2024 architecture, what we kept, what we burned to the ground, and the meeting where two engineers got the same idea at the same time.

DesignRetro
Lia MarchettiPrincipal Engineer
·
34 min read
14 Mar2026
BGP, Anycast, and us: nine months of running the network ourselves.

Why we left a cloud's load balancer, what we got for it, and what we'd recommend you absolutely don't try.

EngineeringNetworking
Yuki TanakaNetwork Engineer
·
26 min read
03 Mar2026
Open-sourcing otel-helpers: the trace context wrapper we use in every service.

Six months in production, 240 services, zero context-loss bugs since. The code is on GitHub. The retrospective is below.

Open sourceObservability
Mira ChenData Platform
·
18 min read

The people writing this stuff.

Forty-one engineers contribute to the blog. These are the most-read this quarter — and the ones who'll happily answer your questions in Discord.

PR
Pavel Ritter
Staff · Build platform
18 posts · 142k reads
SK
Sara Khan
VP Reliability
11 posts · 98k reads
LM
Lia Marchetti
Principal Engineer
9 posts · 86k reads
YT
Yuki Tanaka
Network Engineer
14 posts · 74k reads

Get the next one
in your inbox.

One email, every Monday. Engineering posts, post-mortems, and the occasional design doc we wished we'd published a year ago. No marketing.