Skip to content

DocsRun your app

Scaling

How much compute a replica has, how many replicas an app may run, how many requests a replica takes at once, and where the app runs.

Scaling has three parts: the size of one replica, the number of replicas, and the region. Size and replicas are under Scaling. A change applies with the next deployment.

The size is the CPU and memory of each replica. It is billed per second while the replica runs.

Size CPU Memory
Starter 1 vCPU 512 MB
Small 1 vCPU 1 GB
Medium 2 vCPUs 2 GB
Large 4 vCPUs 4 GB

Start small. If the runtime log says the app ran out of memory, choose the next size.

A replica is one running copy of your app. Replicas sets how many may run at once.

Replicas start when traffic arrives and stop when it goes quiet. With a maximum of 3, a quiet app runs one replica, or none while it sleeps, and a busy one runs up to three.

How many replicas an app may have depends on your plan. See Pricing.

Another replica starts once every running one handles this many requests at the same time. Leave it empty and the size decides:

Size Requests at the same time
Starter 25
Small 25
Medium 50
Large 100

Lower it if each request is heavy. Raise it if requests are light and mostly wait.

Whether an idle app sleeps is set here too, with Keep one replica always on and Wake up fast. See Sleep and wake.

The region is where your app runs. You choose it when you create the app, from regions in North and South America, Europe, Africa, Asia and Oceania. Pick the one closest to your visitors, or to your database.

The region cannot be changed later. To move an app, create a new one in the other region.

Overview shows what each replica is doing right now, and CPU, memory, requests and response time for the last 24 hours.