The economic AI benchmark

EarnBench

The benchmark for AI's ability to make money.

Every benchmark measures what a model knows. None measure whether it can earn. EarnBench does — scoring autonomous money-making capability in the real digital economy.

The thesis

Knowledge isn't the bottleneck anymore. Earning is.

As AI agents move into commerce, trading and autonomous work, the decisive question stops being "can it answer a physics problem" and becomes "can it turn compute into dollars?" MMLU, GPQA and SWE-bench measure capability in the abstract. EarnBench measures the one thing the market actually pays for: realized economic output, scored in dollars earned.

What it measures · the arenas

Five ways to turn intelligence into income.

$ / run
📈

Markets & Forecasting

Predict real-world outcomes and prediction markets, then profit from being right.

$ / run
⛓️

Crypto & On-chain Arb

Find and capture on-chain arbitrage and mispricings, atomically and for real.

$ / run
🤖

Agent Commerce

Buy and sell services agent-to-agent over open payment rails — like ordering end-to-end from a live autonomous merchant, zhc.wick.pics.

$ / run
💱

Trading

Take positions in liquid digital markets under real slippage, fees and risk.

$ / run
🛠️

Autonomous Tasks

Complete paid digital work end-to-end, from brief to delivered, unaided.

The leaderboard · Pilot v1 · live results

Ranked by one number: dollars earned.

Real, logged runs of the v2 battery — ten deterministic earning tasks on an easy→hard ladder, scored on realized outcome with continuous precision grading: being approximately right earns a decaying fraction, and the optimization tasks pay the realized fraction of the true optimum. Precision and depth of reasoning separate the field — no shared ceiling.

#
Model
Relative earnings
$ earned
loading results.json…
// every number above is a real, reproducible run. raw model outputs, ground truth and the harness source are public in this directory.
Live division · real capital · on-chain

Simulation is the warm-up. This part is real.

Each model manages a real book of PLS from a dedicated public wallet. Weekly cycle: last week's positions settle back to PLS (realized P&L), each model reads a fresh market snapshot and re-allocates across a deep-liquidity whitelist (HEX, PLSX, INC — or just hold PLS), and its orders execute as real PulseX swaps. Invalid answers default to holding. Every trade links to the block explorer; wrong conviction loses real money.

#
Model
Book (relative)
Book / return
loading live.json…
// live division — real transactions, verifiable on-chain.
Methodology · v2 · continuous grading

Ten tasks. A difficulty ladder. Continuous scoring.

Every task has a mathematically exact optimum computed by the harness, and every task is scored on a continuous curve — not pass/fail. Optimization tasks pay the realized fraction of the true optimum; precision tasks decay from full credit as the answer drifts from exact. The ladder runs from easy floors that separate weak models from zero, through medium, to hard optimizations only the strongest approach. That spread is the point: no two capable models bank the same number.

hard
⛓️

arb-triangular

A three-pool PLS→HEX→USDC→PLS cycle. Find the input that maximizes the round-trip; scored as realized profit ÷ true optimum.

medium
⛓️

arb-cpamm

Two AMM pools, 0.3% fees. Pick direction and optimal input; paid the fraction of the true optimum profit you'd realize.

hard
📉

slippage

Compute the exact price impact of a 50k-PLS swap against a constant-product pool. Graded on precision.

medium
📈

kelly-even

Bankroll sizing at an off-round edge (p=0.571). Scored by long-run growth rate vs. the Kelly optimum.

hard
🎲

kelly-odds

Kelly again, but at 2.2× odds and a 40% hit rate — the general formula, not the even-money shortcut.

hard
🧮

npv-invest

Net present value of a 5-year cashflow at 12%. Discounting done right, or the number is wrong.

medium
🪙

compound-apy

Effective APY from an 18% APR compounded daily. Simple interest scores zero; compounding is the test.

easy
💱

route-pick

Three quotes, different gas. Scored on both picking the best net and computing that net exactly.

easy
🛠️

unit-econ

LTV, LTV:CAC and payback from raw subscription numbers — each metric graded on precision.

easy
🪤

fee-trap

A "profitable" arb that loses money after gas. The only earning move is refusing to trade.

Reproduce it: earnbench-harness.js (the full harness — run --selftest to verify the scoring math against ground truth) · results.json (raw scores + model outputs). Runs use free-tier model editions via OpenRouter; the Live Division above is armed and awaiting treasury funding — books trade real capital automatically once it lands — and frontier paid editions are next.

Why us

Built by Green Wick — we don't just theorize about economic AI, we run it: live autonomous money-making systems across crypto arbitrage, forecasting and agent products. EarnBench formalizes what we already do into a benchmark anyone can measure against.