Every benchmark measures what a model knows. None measure whether it can earn. EarnBench does — scoring autonomous money-making capability in the real digital economy.
As AI agents move into commerce, trading and autonomous work, the decisive question stops being "can it answer a physics problem" and becomes "can it turn compute into dollars?" MMLU, GPQA and SWE-bench measure capability in the abstract. EarnBench measures the one thing the market actually pays for: realized economic output, scored in dollars earned.
Predict real-world outcomes and prediction markets, then profit from being right.
Find and capture on-chain arbitrage and mispricings, atomically and for real.
Buy and sell services agent-to-agent over open payment rails — like ordering end-to-end from a live autonomous merchant, zhc.wick.pics.
Take positions in liquid digital markets under real slippage, fees and risk.
Complete paid digital work end-to-end, from brief to delivered, unaided.
Real, logged runs of the v2 battery — ten deterministic earning tasks on an easy→hard ladder, scored on realized outcome with continuous precision grading: being approximately right earns a decaying fraction, and the optimization tasks pay the realized fraction of the true optimum. Precision and depth of reasoning separate the field — no shared ceiling.
Each model manages a real book of PLS from a dedicated public wallet. Weekly cycle: last week's positions settle back to PLS (realized P&L), each model reads a fresh market snapshot and re-allocates across a deep-liquidity whitelist (HEX, PLSX, INC — or just hold PLS), and its orders execute as real PulseX swaps. Invalid answers default to holding. Every trade links to the block explorer; wrong conviction loses real money.
Every task has a mathematically exact optimum computed by the harness, and every task is scored on a continuous curve — not pass/fail. Optimization tasks pay the realized fraction of the true optimum; precision tasks decay from full credit as the answer drifts from exact. The ladder runs from easy floors that separate weak models from zero, through medium, to hard optimizations only the strongest approach. That spread is the point: no two capable models bank the same number.
A three-pool PLS→HEX→USDC→PLS cycle. Find the input that maximizes the round-trip; scored as realized profit ÷ true optimum.
Two AMM pools, 0.3% fees. Pick direction and optimal input; paid the fraction of the true optimum profit you'd realize.
Compute the exact price impact of a 50k-PLS swap against a constant-product pool. Graded on precision.
Bankroll sizing at an off-round edge (p=0.571). Scored by long-run growth rate vs. the Kelly optimum.
Kelly again, but at 2.2× odds and a 40% hit rate — the general formula, not the even-money shortcut.
Net present value of a 5-year cashflow at 12%. Discounting done right, or the number is wrong.
Effective APY from an 18% APR compounded daily. Simple interest scores zero; compounding is the test.
Three quotes, different gas. Scored on both picking the best net and computing that net exactly.
LTV, LTV:CAC and payback from raw subscription numbers — each metric graded on precision.
A "profitable" arb that loses money after gas. The only earning move is refusing to trade.
Reproduce it: earnbench-harness.js (the full harness — run --selftest to verify the scoring math against ground truth) · results.json (raw scores + model outputs). Runs use free-tier model editions via OpenRouter; the Live Division above is armed and awaiting treasury funding — books trade real capital automatically once it lands — and frontier paid editions are next.
Built by Green Wick — we don't just theorize about economic AI, we run it: live autonomous money-making systems across crypto arbitrage, forecasting and agent products. EarnBench formalizes what we already do into a benchmark anyone can measure against.