PLATFORM · AUTOSCALE

From zero to 128 pods,
before the spike notices.

Predictive, per-region, per-signal. Scale on RPS, CPU, p95, queue depth — or your own metric. Warm pools keep cold starts at zero milliseconds, and scale-to-zero on staging means you stop paying for idle.

0mscold start, warm pools 128max instances, per region 5built-in scale signals −68%median monthly bill
01Scale signals

Five signals out of the box. One from you.

Mix and match. Combine with weights. Per-service, per-region, per-environment. Or, if your stack is weird, emit your own metric and we'll scale on that.

▸ Default

Requests per second

Window-aware. Smooths over 30s. Scales up in <1s, down with a 90s cooldown.

react time<1s
▸ Compute

CPU & memory

For batch and CPU-bound services. Target utilization with hysteresis to prevent flap-storms.

target70%
▸ Latency

p95 & p99

Scale up the second user-perceived latency starts creeping. Bound by error budget burn-rate.

targetp95 < 80ms
▸ Queue

Workers & backlog

For job runners and Kafka consumers. Scale on lag, queue depth, or oldest-message age.

lag floor< 5s
02Predictive

Scale up before the spike, not after.

We learn your traffic shape — Black Friday, Monday standup, the 14:00 EU rush. Capacity is in place 90 seconds before the curve.

Predicted, not reacted.

Reactive autoscalers see a spike, then chase it — and your users see 503s while pods boot. Scalable trains on 28 days of traffic per service and pre-warms the pool 90s before predicted demand.

  • Lead time+90s
  • Median MAE on next 5 min±3.8%
  • 503s during spike0
  • Training datalast 28d, per service
  • Overridescalable scale --hold
now -30m +30m
actual4.2k rps
predicted11.4k rps
+90s leadscaling now
03Scale to zero

The other half of autoscale.

Staging at 03:00 doesn't need 12 pods. Preview environments don't need anything. Pause when nobody's looking — wake on the first request, in ~140ms.

Before · fixed poolmonthly
checkout-svc · staging · 24/7 × 4 pods$640
api-gateway · staging · 24/7 × 6 pods$960
batch-runner · idle 18h/day × 8 pods$1,420
42 preview environments · always-on$2,180
prod over-provision (peak buffer)$3,300
Total monthly$8,500
After · scale-to-zero + predictivemonthly
checkout-svc · staging · scale-to-zero off-hours$240
api-gateway · staging · scale-to-zero off-hours$360
batch-runner · woken on job submit$280
42 preview environments · pause on inactivity$420
prod predictive · no buffer needed$1,400
Total monthly$2,700

Median customer saves 68% on compute in the first 30 days — same throughput, same latency.

"Black Friday hit and we didn't get a single page. Capacity was just already there."
SR
Sofia ReyesPlatform lead · Cumulus.bank · 4× peak handled, zero alerts
Peak RPS handled1.2k → 48k
Pages during Black Fridayzero.
Monthly compute spend71%

Stop paying for idle.
Stop running out of headroom.

Plug your repo in, and we'll learn your traffic shape in 28 days. Median customer cuts compute spend by 68% — without a single perf regression.