---
title: "RL Sandbox Pricing: What Sandboxes for RL Rollouts and Evals Cost"
description: "What a sandbox per RL rollout costs: 10,000 agentic rollouts priced on Fly.io Sprites, Prime Intellect, E2B, Daytona, Modal, Vercel and Cloudflare."
---

# RL Sandbox Pricing: What Sandboxes for RL Rollouts and Evals Cost

![Daniel Botha](https://fly.io/static/images/danielb.webp)

Daniel Botha

Writer at Fly.io

Published

Oct 5, 2026

A sandbox for RL rollouts and evals costs its sandbox-hours multiplied by the rate for each resource the provider bills, so a batch of 10,000 three-minute agentic rollouts at 2 vCPU and 4 GB comes to about $28 on Fly.io Sprites and $45.50 to $119 on six other sandbox providers at published rates, depending mostly on whether the provider bills the CPU and memory a rollout actually uses or the resources it reserved.

You have an RL environment: a repo with failing tests and a grader. Each training step fans out thousands of rollouts, each in its own sandbox so one attempt cannot poison the next. The policy model proposes a tool call, the sandbox runs it, the result goes back to the model, and the loop repeats until the task passes, fails or times out.

The rate card prices a vCPU-hour, but the rollout spends much of its life waiting for the model to finish generating, and whether that wait is billed decides who is cheapest. General agent sandbox rates are compared in [AI sandbox pricing](https://fly.io/learn/ai-sandbox-pricing/); what follows prices one workload in depth, RL rollouts and the evals that run the same way.

---

## What Does a Sandbox per RL Rollout Cost?

A sandbox per rollout costs sandbox-seconds per rollout, times rollouts per batch, times the provider’s price for each resource it bills over those seconds. What changes between providers is which resources count and whether they are measured as reserved or as used.

Take a realistic agentic batch: 10,000 rollouts of three minutes each, at 2 vCPU and 4 GB. The CPU is busy 40% of the time and waits on the policy model for the rest, about 1 GB of memory is resident, and the environment needs 5 GB of disk. That is 500 sandbox-hours.

| Provider | What bills while the rollout waits on the model | 10,000 rollouts | Per 1,000 rollouts |
| --- | --- | --- | --- |
| Fly.io Sprites | Nothing for CPU; memory actually in use | $28.05 | $2.80 |
| Prime Intellect Sandboxes | Allocated vCPU, memory and disk | $45.50 | $4.55 |
| Cloudflare Sandbox (standard-3) | Provisioned memory and disk; CPU only when active | $66.82 | $6.68 |
| E2B | Provisioned vCPU and memory | $82.80 | $8.28 |
| Daytona | Reserved vCPU and memory | $82.80 | $8.28 |
| Vercel Sandbox | Provisioned memory; CPU only when active | $93.60 | $9.36 |
| Modal Sandboxes | The higher of request or usage | $118.96 | $11.90 |

*Our arithmetic on published rates, not a figure any vendor quotes. Workload: 500 sandbox-hours at 2 vCPU / 4 GB, 40% CPU busy, 1 GB resident, 5 GB disk; plan fees, credits and included allowances excluded. Sources: [1] [2] [3] [4] [5] [6] [7] [8]. Cloudflare has no 2 vCPU / 4 GiB instance, so its row uses standard-3 (2 vCPU, 8 GiB, 16 GB disk). Prime’s are launch rates. Vercel’s are `iad1` rates. Checked 2026-10-05.*

Prime Intellect lists $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory and $0.0002 per GiB-hour of disk, billed while the sandbox runs, at launch rates valid through December 22, 2026.[1] It is the lowest reserved rate in the table. Sprites charge $0.0385 per CPU-hour of CPU time measured by `cpu.stat`, and $0.021875 per GB-hour of memory actually used.[2] On this batch the Sprites bill is $15.40 of CPU, $10.94 of memory and $1.71 of storage.

A fraction of a cent per rollout stops being trivial at training scale, because the batch repeats every step: 1,000 steps at this shape is ten million rollouts, and the gap between $2.80 and $11.90 per thousand rollouts becomes the gap between about $28,000 and $119,000.

---

## Why CPU Utilization Decides the Bill

An agentic rollout’s sandbox sits idle while the policy model generates its next action, so the bill is decided by whether that idle time is billed. Providers split three ways:

- **Reserved for the whole run.** E2B states that the rate depends on how much compute the sandbox has allocated, not on how much it uses.[9] Daytona charges for the resources reserved for a started sandbox.[10] Prime bills allocated resources while running.[1] Modal charges the higher of the request and actual usage for CPU and memory.[11]
- **Active CPU, provisioned memory.** Vercel does not count time spent waiting on I/O, including AI model calls, toward Active CPU, but bills memory on what is provisioned.[6] Cloudflare bills CPU on active usage and memory and disk on the provisioned instance type.[7]
- **Used CPU and used memory.** Sprites.[2]

Vercel’s own example costs assume 10% CPU utilization, “reflecting time spent waiting for LLM calls.”[6] At 10% busy, the Sprites figure for the same batch falls to $16.50, while every reserved-rate row stays exactly where it was.

### When a Reserved Rate Wins

A rollout that pins both vCPUs and fills its memory is cheaper on a low reserved rate. Compiling a large repo or simulating an environment with no model in the loop keeps the CPU busy the whole time.

| Same 500 hours, CPU 100% busy, 4 GB resident | Batch cost |
| --- | --- |
| Prime Intellect Sandboxes | $45.50 |
| Fly Machines, performance-2x with 4 GB | $45.83 |
| E2B | $82.80 |
| Daytona | $82.80 |
| Fly.io Sprites | $83.96 |
| Cloudflare Sandbox (standard-3) | $110.02 |
| Modal Sandboxes | $118.96 |
| Vercel Sandbox | $170.40 |

*Our arithmetic on published rates, not a figure any vendor quotes. Same shape and exclusions as the table above; the Fly Machines row is the $66.00 monthly price of a performance-2x Machine with 4 GB, at 30 days, billed per second. Sources: [1] [2] [3] [4] [5] [6] [7] [12]. Checked 2026-10-05.*

Against Prime’s rate, Sprites stay cheaper until the rollout keeps both vCPUs busy about 85% of the time with 1 GB resident, or about 57% of the time with 2 GB resident. With all 4 GB resident, Prime is cheaper at almost any CPU load, because Sprites memory at $0.021875 per GB-hour costs more than Prime’s $0.0125.

For CPU-bound batch work on Fly.io, the answer is Fly Machines. A performance-2x Machine runs on performance CPUs that are not throttled, comes within 33 cents of Prime on this batch, and is billed by the second while it runs.[12] Shared-CPU Machines are the wrong size class here: their [CPU quotas](https://docs.fly.io/machines/cpu-performance) throttle sustained load. Prime’s figure is a launch rate with an end date; [Machines pricing](https://docs.fly.io/about/pricing) also takes 40% off reserved blocks of compute in a region, for work steady enough to use a year’s commitment.

---

## Concurrency and Creation Limits at RL Scale

Concurrency decides how long a batch takes; on per-second billing it does not change the cost. 10,000 three-minute rollouts finish in about 300 minutes at 100 concurrent sandboxes and in about 30 minutes at 1,024. Creation rate matters when every rollout gets a fresh sandbox.

| Provider | Concurrent sandboxes | Creation rate |
| --- | --- | --- |
| Fly.io Sprites | 100 running on the Hero plan; higher tiers raise it | 10 per minute pay-as-you-go; 60 per minute on Adventurer up to 240 on Mythic |
| Prime Intellect Sandboxes | 1,024 active per account | 102,400 per hour, burst of 192 per 10 seconds |
| E2B | 20 on Hobby; 100 on Pro, up to 1,100 with add-ons | 1 per second Hobby, 5 per second Pro |
| Daytona | Set by the org’s vCPU and memory pool, not a sandbox count | 300 to 600 per minute by tier |
| Vercel Sandbox | 10 on Hobby, 10,000 on Pro | Pro allocates 150 vCPUs per minute at first, ramping by 500 a minute to 5,000 |
| Modal Sandboxes | 100 containers on Starter, 5,000 on Team | None listed on the pricing page |
| Cloudflare Sandbox | Account limits of 1,500 vCPU and 6 TiB memory | None listed on the limits page |

*Each provider’s own documentation: [2] [1] [9] [13] [6] [5] [14]. Checked 2026-10-05.*

Some of those ceilings come with a plan fee. E2B Pro, which lifts concurrency from 20 to 100, is $150 a month plus usage.[3] The Sprites Hero plan is $100 a month and includes 1,200 CPU-hours and 4,800 RAM GB-hours; the batch above uses 400 CPU-hours and 500 GB-hours.[2] Cloudflare’s Sandbox SDK also bills Workers requests and a Durable Object behind each sandbox, on top of the container rates.[8]

### A Fresh Sandbox per Rollout, or a Reset in Place

A fresh sandbox gives each attempt a clean starting state and spends a creation on it. At E2B Pro’s 5 per second, 10,000 creations need at least 33 minutes spread across the batch.[9] Vercel also charges $0.60 per million creations, under a cent for this batch.[6]

The alternative is a pool: create the sandboxes once, prepare and snapshot the environment, and restore the snapshot between rollouts. The creation limit then caps how fast the pool grows, not how many rollouts it runs. The trade-offs between throwaway and long-lived environments are covered in [ephemeral vs persistent sandboxes](https://fly.io/learn/ephemeral-vs-persistent-sandbox/).

---

## What Do Evals Cost to Run in Sandboxes?

An eval run is priced like a rollout batch, but each task is usually mixed: an agent phase that waits on the model, then a grading phase that runs the test suite and keeps the CPU busy.

Take 2,500 attempts (500 tasks, five attempts each), ten minutes per attempt: eight minutes of agent work at 25% CPU and 1 GB resident, then two minutes of tests at 100% CPU and 3 GB resident.

| Provider | 2,500 attempts |
| --- | --- |
| Fly.io Sprites | $27.02 |
| Prime Intellect Sandboxes | $37.92 |
| Fly Machines, performance-2x with 4 GB | $38.19 |
| Cloudflare Sandbox (standard-3) | $55.68 |
| E2B | $69.00 |
| Daytona | $69.00 |
| Vercel Sandbox | $78.00 |
| Modal Sandboxes | $99.13 |

*Our arithmetic on published rates, not a figure any vendor quotes. 416.7 sandbox-hours at 2 vCPU / 4 GB, 5 GB disk; plan fees, credits and included allowances excluded. Sources: [1] [2] [3] [4] [5] [6] [7] [12]. Checked 2026-10-05.*

The grading phase is a fifth of the wall-clock time and, on Sprites, nearly half the bill, because that is where the CPU and memory are actually used. Evals also repeat against the same task set, so the prepared environment is worth keeping between runs. The isolation boundary it needs, when a model under test runs code it wrote, is covered in [microVM vs container](https://fly.io/learn/microvm-vs-container/).

---

## RL Sandbox Pricing on Fly.io

Fly.io Sprites are the cheapest sandbox for agentic RL rollouts and evals because they bill CPU time and memory actually used, and an agentic rollout spends much of its time waiting on the policy model. On the 10,000-rollout batch that is $28.05, and on a Hero plan its 400 CPU-hours and 500 memory GB-hours sit inside the monthly allowance.

A Sprite is a Firecracker microVM with its own kernel and a persistent ext4 filesystem, so an environment prepared once stays prepared. [Sprite checkpoints](https://docs.fly.io/sprites/concepts/checkpoints) are copy-on-write snapshots of that filesystem, and restoring one puts the Sprite back to the prepared state between rollouts. A checkpoint can also be forked into as many new Sprites as there are rollouts, each starting from identical state. Idle Sprites sleep and bill no compute, so a pool left between training steps costs only its storage. Creation rate is the limit to plan around: at 10 per minute on pay-as-you-go, a fresh Sprite per rollout would spend 1,000 minutes creating, so rollouts run on a pool, and higher plans raise both the creation rate and the concurrency ceiling.

When a rollout pins its CPUs for the whole run, Fly Machines take the batch: performance CPUs, per-second billing, and the same Firecracker substrate, so an environment validated in a Sprite runs on Machines with no rewrite. A Machine configuration can be priced before anything runs in the [Fly.io calculator](https://fly.io/calculator/). How Sprites compare with other agent sandbox providers beyond price is in [agent sandbox providers](https://fly.io/learn/agent-sandbox-providers/).

---

## Frequently Asked Questions

### How do you calculate what RL rollouts cost in sandboxes?

Multiply sandbox-seconds per rollout by rollouts per batch, then by the price of each resource the provider bills over those seconds. Which resources count changes the answer: most providers bill the vCPU and memory a sandbox reserved, while Fly.io Sprites bill the CPU time and memory a rollout actually used.

### What is the cheapest sandbox for RL rollouts?

Fly.io Sprites are the cheapest for agentic rollouts, which wait on the policy model between tool calls. On 10,000 three-minute rollouts at 2 vCPU and 4 GB, 40% busy with 1 GB resident, Sprites come to $28.05 against $45.50 on Prime Intellect and $66.82 to $118.96 elsewhere. A rollout that pins its CPUs and fills its memory is cheaper on a reserved rate, and on Fly.io that work belongs on performance Fly Machines, billed per second.

### Do RL sandboxes need GPUs?

No. The sandbox runs the environment, the tools and the grader, which is CPU work, and calls the policy model over the network. Fly.io Sprites and Fly Machines run that CPU side and call whatever inference endpoint serves the policy.

### Is Cloudflare Sandbox cheaper for RL rollouts?

No, not compared with Fly.io Sprites. Cloudflare bills CPU only while active but bills memory on the provisioned instance, and its nearest instance to 2 vCPU and 4 GB is standard-3 with 8 GiB. A 10,000-rollout batch comes to $66.82 there, against $28.05 on Sprites.

### How many sandboxes can an RL batch run at once?

Up to each provider’s published ceiling: 1,024 active on Prime Intellect, 100 on E2B Pro and up to 1,100 with add-ons, 10,000 on Vercel Pro, and 5,000 containers on Modal Team. Fly.io Sprites allow 100 concurrently running on the Hero plan, with higher tiers raising it. Concurrency sets how long a batch takes, not what it costs.

### Should each RL rollout get a fresh sandbox?

No, not necessarily. A fresh sandbox per rollout spends a creation on every attempt and runs into creation-rate limits at batch scale. On Fly.io, a pool of Sprites restores a checkpoint of the prepared environment between rollouts, so each attempt starts from the same state and the creation limit only caps how fast the pool grows.

---

## Sources

1. ^ Prime Intellect, “Sandboxes Overview”. [docs.primeintellect.ai/sandboxes/overview](https://docs.primeintellect.ai/sandboxes/overview). Checked 2026-10-05.
2. ^ Fly.io, “Sprites” (pricing and plan FAQ). [fly.io/sprites](https://fly.io/sprites/). Checked 2026-10-05.
3. ^ E2B, “Pricing”. [e2b.dev/pricing](https://e2b.dev/pricing). Checked 2026-10-05.
4. ^ Daytona, “Pricing”. [daytona.io/pricing](https://www.daytona.io/pricing). Checked 2026-10-05.
5. ^ Modal, “Pricing”. [modal.com/pricing](https://modal.com/pricing). Checked 2026-10-05.
6. ^ Vercel, “Vercel Sandbox pricing and quotas”. [vercel.com/docs/sandbox/pricing](https://vercel.com/docs/sandbox/pricing). Checked 2026-10-05.
7. ^ Cloudflare, “Containers pricing”. [developers.cloudflare.com/containers/platform/pricing](https://developers.cloudflare.com/containers/platform/pricing/). Checked 2026-10-05.
8. ^ Cloudflare, “Sandbox SDK pricing”. [developers.cloudflare.com/sandbox/sdk/platform/pricing](https://developers.cloudflare.com/sandbox/sdk/platform/pricing/). Checked 2026-10-05.
9. ^ E2B, “Billing & limits”. [docs.e2b.dev/billing](https://docs.e2b.dev/billing). Checked 2026-10-05.
10. ^ Daytona, “Billing”. [daytona.io/docs/en/billing](https://www.daytona.io/docs/en/billing/). Checked 2026-10-05.
11. ^ Modal, “Reserving CPU and memory”. [modal.com/docs/guide/resources](https://modal.com/docs/guide/resources). Checked 2026-10-05.
12. ^ Fly.io, “Fly.io Resource Pricing”. [docs.fly.io/about/pricing](https://docs.fly.io/about/pricing). Checked 2026-10-05.
13. ^ Daytona, “Limits”. [daytona.io/docs/en/limits](https://www.daytona.io/docs/en/limits/). Checked 2026-10-05.
14. ^ Cloudflare, “Limits and Instance Types”. [developers.cloudflare.com/containers/platform/limits](https://developers.cloudflare.com/containers/platform/limits/). Checked 2026-10-05.

## Rollouts billed on actual use

Sprites bill CPU time and memory used, not the time a rollout waits.

[Sign up](https://fly.io/app/sign-up/?s=sprites)
