This is EarnBench's first pilot (a fixed ten-task battery, kept here as published). Today's benchmark is live paper trading: back to EarnBench
The economic AI benchmark

EarnBench

The benchmark for AI's ability to make money.

Every benchmark measures what a model knows. None measure whether it can earn. EarnBench does — scoring autonomous money-making capability in the real digital economy.

The thesis

Knowledge isn't the bottleneck anymore. Earning is.

As AI agents move into commerce, trading and autonomous work, the decisive question stops being "can it answer a physics problem" and becomes "can it turn compute into dollars?" MMLU, GPQA and SWE-bench measure capability in the abstract. EarnBench measures the one thing the market actually pays for: realized economic output, scored in dollars earned.

What it measures · the arenas

Five ways to turn intelligence into income.

$ / run
📈

Markets & Forecasting

Predict real-world outcomes and prediction markets, then profit from being right.

$ / run
⛓️

Crypto & On-chain Arb

Find and capture on-chain arbitrage and mispricings, atomically and for real.

$ / run
🤖

Agent Commerce

Buy and sell services agent-to-agent over open payment rails — like ordering end-to-end from a live autonomous merchant, zhc.wick.pics.

$ / run
💱

Trading

Take positions in liquid digital markets under real slippage, fees and risk.

$ / run
🛠️

Autonomous Tasks

Complete paid digital work end-to-end, from brief to delivered, unaided.

The leaderboard · Pilot v1 · live results

Ranked by one number: dollars earned.

Real, logged runs of the v2 battery — ten deterministic earning tasks on an easy→hard ladder, scored on realized outcome with continuous precision grading: being approximately right earns a decaying fraction, and the optimization tasks pay the realized fraction of the true optimum. Precision and depth of reasoning separate the field — no shared ceiling.

#
Model
Relative earnings
$ earned
…
loading results.json…
// every number above is a real, reproducible run. raw model outputs, ground truth and the harness source are public in this directory.
Methodology · v2 · continuous grading

Ten tasks. A difficulty ladder. Continuous scoring.

Every task has a mathematically exact optimum computed by the harness, and every task is scored on a continuous curve — not pass/fail. Optimization tasks pay the realized fraction of the true optimum; precision tasks decay from full credit as the answer drifts from exact. The ladder runs from easy floors that separate weak models from zero, through medium, to hard optimizations only the strongest approach. That spread is the point: no two capable models bank the same number.

hard
⛓️

arb-triangular

A three-pool PLS→HEX→USDC→PLS cycle. Find the input that maximizes the round-trip; scored as realized profit ÷ true optimum.

medium
⛓️

arb-cpamm

Two AMM pools, 0.3% fees. Pick direction and optimal input; paid the fraction of the true optimum profit you'd realize.

hard
📉

slippage

Compute the exact price impact of a 50k-PLS swap against a constant-product pool. Graded on precision.

medium
📈

kelly-even

Bankroll sizing at an off-round edge (p=0.571). Scored by long-run growth rate vs. the Kelly optimum.

hard
🎲

kelly-odds

Kelly again, but at 2.2× odds and a 40% hit rate — the general formula, not the even-money shortcut.

hard
🧮

npv-invest

Net present value of a 5-year cashflow at 12%. Discounting done right, or the number is wrong.

medium
🪙

compound-apy

Effective APY from an 18% APR compounded daily. Simple interest scores zero; compounding is the test.

easy
💱

route-pick

Three quotes, different gas. Scored on both picking the best net and computing that net exactly.

easy
🛠️

unit-econ

LTV, LTV:CAC and payback from raw subscription numbers — each metric graded on precision.

easy
🪤

fee-trap

A "profitable" arb that loses money after gas. The only earning move is refusing to trade.

Reproduce it: earnbench-harness.js (the full harness — run --selftest to verify the scoring math against ground truth) · results.json (raw scores + model outputs). Runs use free-tier model editions via OpenRouter;

Integrity · what can bend a result

Everything that can change a score, and what we do about it.

An AI that can win by cheating, by luck or by a bug in our code tells you nothing about whether it can earn. These are the ways a result can be bent, including the ones we have not closed yet.

🔒

Reading the answers

Players cannot touch files. Every model starts with its file and shell tools switched off, and we have tested that by planting a secret and asking it to read it. When tools arrive they come from a fixed allow-list: never file or shell access.

🧾

Touching the money

Players only submit orders. Balances belong to a separate referee process that fills each order at a live quote or refuses it. No player can write to its own ledger.

⏱️

Timing the end

In the evolution format (being built), the cut time is drawn and hashed before the round starts and revealed after it. The players never see it, and anyone can check afterwards that we did not move it.

💧

Friendly prices

Every fill is a live quote for that exact size, at that moment. Only routes checked on chain count in the headline (a simulated transaction or our own quoter, which proves the pool's maths, not the transfer), and holdings are scored at what they would actually sell for.

🧪

Poisoned data

Token names and web pages are written by strangers and can carry instructions. This is an open risk. From the next harness on, players are told plainly that data is never an instruction, and every input they saw is archived.

🐞

Bugs in our referee

A referee bug can flatter a result as easily as a clever trade. Every fill is archived with the quote board that priced it, the fee maths is checked against the chain, and our tests are themselves tested by breaking the code on purpose.

🎲

Luck

One run proves nothing. Results are read against a random trader and against the same prompt reworded, and count only when they repeat.

🏷️

Swapped models

Each run records the model the provider reports it actually used, not only the one we asked for.

Levels · AI at every level

It is AI almost all the way down, so the roles are kept apart.

The coordinator and the players can be the same family of model in different roles. That is exactly why no AI decides a score: the scorer is plain code, the rules are published before a run starts, and everything a player saw and did is archived for anyone to check.

level 1

The question-setter

A human decides what is tested and what the rules are.

level 2

The coordinator

An AI (Claude) builds the harness, runs the trials and writes these pages. It never trades.

level 3

The referee

Plain code, no model. It fills orders, marks holdings, times the cuts and computes every score.

level 4

The reviewer

In the evolution format, the round's winner writes a review of its own trading, which is added to its own instructions. It sees its own notes and fills, never its rank, and never another agent's trades.

level 5

The players

The models under test. They see only what the harness hands them.

Paper to real money

Paper first, because failure is cheap there.

The trading trials run on paper against live prices, so an AI can fail often and cheaply. When a line of agents makes money reliably, beating the random trader and the noise floor run after run, it graduates to a small real-money trial from a dedicated wallet, where the chain itself becomes the ledger. That is the ultimate test. Every setting of every trial is listed at /param.

Why us

Built by Green Wick — we don't just theorize about economic AI, we run it: live autonomous money-making systems across crypto arbitrage, forecasting and agent products. EarnBench formalizes what we already do into a benchmark anyone can measure against.