Deep dives, post-mortems, design docs we wish we'd written before we shipped, and the occasional unhinged opinion. No PR-reviewed thought leadership. If we got it wrong, we say so.
What our metrics said was one cascading failure. What we found after going through five days of trace data, log streams, and one extremely confused engineer.
A 200-line model in Rust beats the reactive autoscaler we shipped two years ago on every metric we care about — including cold starts during the EU rush hour.
The design doc behind v4. How we got rollbacks to 4 seconds, regional fan-out to "free," and convinced our security team that immutable images are not a stretch goal.
PgBouncer, PgCat, Supavisor, and 11 others — same workload, same hardware, 36 hours. The results were not what we expected, and one was disqualified for cheating.
We migrated from Kafka. Three things we miss, four things we don't, one thing we still can't believe works as well as it does.
Just routing. Just BGP. Just patience. A diary of the 6 weeks we spent at three IXPs in Frankfurt, Marseille, and Singapore — and what came out of it.
The bill, the people, the on-call rotation. Not the marketing version — the engineering version, with line items.
A retrospective on the 2024 architecture, what we kept, what we burned to the ground, and the meeting where two engineers got the same idea at the same time.
Why we left a cloud's load balancer, what we got for it, and what we'd recommend you absolutely don't try.
Six months in production, 240 services, zero context-loss bugs since. The code is on GitHub. The retrospective is below.
One email, every Monday. Engineering posts, post-mortems, and the occasional design doc we wished we'd published a year ago. No marketing.