---
title: "Autoscaling"
description: "Autoscaling on Fly.io. Machines wake on the first request and stop when traffic goes quiet, and you pay nothing for CPU or memory while they are stopped."
---

# *Autoscaling* on Fly.io

Why pay for servers that are just sitting there eating Cheetos? Your app should scale up when it's busy and scale down when it's not. Autoscaling on Fly.io means you only run what you need, when you need it. It's your accountant's favorite feature.

[Get Started](https://fly.io/app/sign-up)

![Autostop and Autostart](https://fly.io/phx/ui/images/autoscaling_1-0f7d0a3a77c07609289b9b7b6b0c3836.png?vsn=d)

## Autostop/Autostart: The "Only Pay for What You Use" Button

A user calls your app, your Machines wake up. Like ... really, really fast. No traffic? Machines go back to sleep. No CPU or RAM charges while stopped. It's easy to configure this behavior to cover exactly what your project needs.

- Wake up in milliseconds when requests arrive
- Stop automatically during idle periods
- Zero CPU/RAM charges when stopped

- Set minimum Machines to keep warm
- Perfect for sporadic traffic patterns
- Enabled by default for new apps

## Metrics-Based Autoscaling: Because Queue Depth > Request Count

Sometimes "number of HTTP requests" isn't the right metric. Got background workers chomping through a job queue? Temporal workflows piling up? Scale based on what actually matters to your app: queue depth, pending work, custom metrics from Prometheus, whatever keeps you up at night.

Scale based on queue depth, pending jobs, or custom metrics

Pull metrics from Prometheus or Temporal

Write scaling rules with expressions and arithmetic

Create or destroy Machines dynamically

Scale multiple apps with common naming patterns

![Metrics-Based Autoscaling](https://fly.io/phx/ui/images/autoscaling_2-131df9f1c9c8d942bb0c9d6dc5bcf4f7.png?vsn=d)

## What Even is This Magic?

Fly Proxy sits at the edge and watches traffic. When a request shows up for a sleeping Machine, the proxy wakes it up faster than you can say "cold start problem." Machines boot in milliseconds, handle the request, and go back to sleep when things quiet down. You configure the behavior, we handle the orchestration.

Fly Proxy detects incoming traffic and wakes Machines instantly

Configure stop/start behavior in your fly.toml

Set minimum Machines running to avoid cold starts

Machines only get charged when they're actually running

[Read the autoscaling docs](https://docs.fly.io/reference/autoscaling)

## Read the Docs

- [Autoscaling Overview](https://docs.fly.io/reference/autoscaling): Learn about the different autoscaling strategies on Fly.io

- [Autostop/Autostart](https://docs.fly.io/launch/autostop-autostart): Configure automatic Machine stopping and starting based on traffic

- [Autoscale by Metric](https://docs.fly.io/launch/autoscale-by-metric): Scale based on custom metrics like queue depth or pending work

- [Configuration Reference](https://docs.fly.io/reference/configuration): See all the autoscaling settings in fly.toml

## Stop Paying for Idle Servers

![Get Started with Autoscaling](https://fly.io/phx/ui/images/autoscaling_3-33a2f316fa16f6f5cde320044ae696b7.png?vsn=d)

Whatever happened to the promise of paying only for what you use in cloud infra? With Fly.io autoscaling, that's actually true. Machines wake up when there's work to do and go back to sleep when there isn't. No babysitting required.

[Try It Free](https://fly.io/app/sign-up)
