The live race
for cloud sandboxes.
- Pick a workload.
- ABTwo sandboxes line up on the grid and run side by side —
- Streaming logs, test results, real $/run.
- Lights out — the metrics decide who takes the flag.
Live timing
Four sandboxes on a hot lap — anonymized as A and B in the real arena. Cold-start wall time, lowest is leading. Hover the track to read each car's livery.
lights out · A vs B
- P101E2B277 ms
- P202Blaxel575 ms
- P303Declaw1.05 s
- P404Vercel Sandbox1.23 s
- P505Daytona1.59 s
- P606CodeSandbox2.93 s
- P707Dedalus5.00 s
- P808Sprites6.82 s
- +6 sandboxes awaiting first run
277 ms
cold start
$0.00002
per run
14
on the grid
9
audited
- 01Line up on the grid
Pick a workload. Two sandboxes take their slots, anonymized as A and B.
- 02Identical conditions, blind
Same bytes, same vCPU and RAM, same region — declared on every match. Then: lights out.
- 03Watch them race
Run lines stream live, phases tick green, the HUD fills with provision, exec and $ per run.
- 04Take the flag
Objective metrics decide the winner; you vote which you'd ship, then identities are revealed.
Thirteen contenders, mapped
Stateless ↔ stateful on one axis, lightweight ↔ compute-heavy on the other — every contender mapped on the same board, no house favorite.
One pit wall, every signal
One workload goes in; identical runs come back out as streaming logs, phase metrics and real $/run — wired into a single race the moment you press go.
Thirteen on the grid
Each contender implements one adapter contract — `create / exec / destroy / estimateCost` — so every sandbox laps the same circuit with no per-provider special cases.
Auditable, not just subjective
A run booted, executed and passed — or it didn't, in a measurable number of milliseconds and dollars.
| Property | LMArena | Design Arena | Sandbox Arena |
|---|---|---|---|
| Contestants | LLMs | UI generations | Cloud sandboxes |
| Artifact | Two answers | Two rendered UIs | Two live runs (logs + tests) |
| Ranking | Crowd Elo | Crowd Elo | Objective composite + crowd Elo |
| Objectivity | Subjective | Subjective | Mostly objective, vote as tiebreak |
| Cost shown | No | No | Yes ($/run, head to head) |
| Reproducible | Partially | No | Yes (open workloads + raw results) |
Credibility, not vibes
An arena with a house favorite is worthless. Every number is reproducible from open inputs — here is the paperwork.
Reproducible
Workloads, scoring and profiles are in the repo. `pnpm check:scores` regenerates the exact leaderboard numbers.
Blind pairing
Contenders run anonymized as Sandbox A / B and stay hidden until after you vote, so brand never biases the result.
No house favorite
No provider is special-cased — each drops the workloads it should, the filesystem storm and the long build included. Wins are only believable when the losses are visible.
A different approach to benchmarking
Credibility is the product
Workload definitions, the scoring formula and every performance profile live in the repo. If a skeptic cannot reproduce a number from the open catalog and raw results, the number does not ship.
Who's on pole right now
| Pos | Provider | Objective | Latency p50 / p99 | Crowd Elo | Record (W–L–T) | Runs | |
|---|---|---|---|---|---|---|---|
| 01 | 98.9 | 1.18 s / 1.84 s | 952 | 36–5–1 | 42 | ||
| 02 | 94.3 | 1.83 s / 3.95 s | 1316 | 36–17–1 | 52 | ||
| 03 | 92.2 | 2.54 s / 25.91 s | 1107 | 47–32–2 | 81 | ||
| 04 | 88.1 | 1.01 s / 11.14 s | 1330 | 116–26–1 | 143 | ||
| 05 | 79.9 | 5.09 s / 14.13 s | 878 | 24–18–1 | 43 | ||
| 06 | 69.1 | 3.05 s / 17.46 s | 1176 | 61–40–2 | 103 |
See what the arena can do
Built for ambitious infra. Powered by open, reproducible benchmarks. Engineered to get the numbers right.
Start a race