Questions and answers

Everything you might ask.

What EarnBench does, how it works, what people tend to misread, and every limit we know of, with what we are doing about it. 45 questions, short answers.

What EarnBench does wellHow it worksCommon misreadingsLimits, and what we are doing about themThe bigger picture

What EarnBench does well

What is EarnBench?

A benchmark for one question: can an AI trading setup make money in a real market, by more than luck? We test the whole stack (model, agent harness, prompt and tools), not the model alone.

Are the prices real?

Yes. The money is paper, but every price is live. Each order fills at a live quote for exactly its size, taken at that moment (our own route first, rival aggregators when ours has no price), so the market can't be replayed or memorised.

Can a fill look better than it would really be?

We work hard to stop that. The route's own transaction is simulated on the latest block, so token taxes and transactions that would fail are charged as they would be on chain, and the headline score counts only fills on routes checked on chain. A second line, which counts every fill, is published beside it.

Can a model mark its own homework?

No. A model never holds a balance. A separate referee keeps the books, fills orders and values holdings; the model can only send orders.

How is a result scored?

By liquidation value: every holding is valued with a live, full-size sell quote, not the mid price. That is what the bot could actually walk away with. Holding ETH scores exactly zero.

How do you stop a lucky run looking like skill?

Every run sits beside a random trader in the same window. We also run copies of one model side by side to measure how far luck alone moves a score. A difference only counts once it clears that noise, and a single 1-hour run is treated as an anecdote, not a result.

Can the setup be quietly changed after the fact?

Every harness and prompt is identified by a SHA-256 hash, recorded with each run. Registered tests are frozen by those hashes before they start, and any change is written down as a dated deviation. The methods paper keeps a dated revision log.

Why an entrance exam?

So that a bad result means a bad decision, not a broken answer. Before it trades, a model must play three real trading turns without a single error: a valid answer within the limits, real tool use, and orders that are affordable and use only addresses it was shown. Results are on the entrance exam page.

Do costs count?

Yes. Gas, pool fees, token taxes and slippage are all paid. Each run is also reported after compute: the trading result minus what the model's own calls would cost at its paid list price.

How is EarnBench different from AI trading contests?

Most contests give each model the same prompt and rank the returns. We test the whole setup (model, harness, prompt and tools), a code referee rather than the model fills and checks every trade, comparisons are made within one window against a random trader and copies of the same model, tests are registered before they run, and results are shown after the cost of the model's own calls. The methods paper compares us with the closest projects.

Why not have an AI judge the quality of the reasoning?

Because the question is whether it makes money. A judge model can be persuaded by a confident explanation of a losing trade; a referee that only counts what the wallet would really hold cannot.

Could a model have memorised the market?

No. It trades prices that did not exist when it was trained, so there is no past answer to remember. That is a common weakness of backtests on historical data, and a live market avoids it.

Is anything hidden?

No. Every run has a public page with each turn, the model's written plan, every order and fill, and the reason for every refusal. Every setting is on the parameters page and every rule on the rules page.

How it works

What market do the bots trade?

Robinhood Chain (chain id 4663), a public EVM network with Uniswap-style pools and many newly launched tokens. EarnBench is not affiliated with Robinhood.

How much does each bot start with?

US$100 of ETH, priced when its run opens.

How often does a bot act?

It is woken on a fixed schedule (every 10 minutes by default). Each time it sees its data and its own book and may send up to five orders: buy, sell, or open or close a liquidity position.

What can a model see and use?

Whatever its harness gives it, and nothing else. Each harness's exact tools and limits are on the tools page.

Does a bot get time to prepare?

Yes. Before its window opens it gets one research turn of up to 20 minutes to write a trading plan. The plan is published with the run, so every later trade can be judged against it.

How long are runs?

1-hour runs prove a model can trade end to end and let us try variations quickly. Longer runs, such as 4 hours, are used for comparisons.

What if a bot does nothing?

In current trials, a bot is told that taking no action disqualifies it. A run that then places no trade at all is left out of the results.

Which models are tested?

Models we can run without per-call charges: free tiers from several providers, and Claude on a subscription. The entrance exam page lists every model that has sat the exam and the stats page counts them.

What is Space Bunny Alpha?

An anonymous preview model on OpenRouter, free while its lab keeps it unnamed. If the lab later confirms it was the same model all along, its results will be renamed to the real model. If the released model is different, they stay under the preview name.

Common misreadings

Is real money traded?

No. Every trade is on paper, priced and checked against the live market.

Does a top score mean I should trade with that model?

No. It means that setup did well in those windows on this market. Past windows promise nothing about future ones, and paper fills are kinder than real ones in two known ways (see Limits).

Why does a results page show a model's best run?

The entrance exam and stats pages show each model's best run, always with how many runs it had and their average beside it. Registered tests use every run, not the best one.

Why is the score in ETH and not dollars?

Every bot in a window sees the same ETH price move, so the ranking is the same either way. In ETH, doing nothing scores exactly zero, which makes the baseline obvious.

Do the bots trade against each other?

No. They share a market, not an order book. A paper trade does not move the pool, so one bot cannot change another's prices.

Doesn't the smartest model on the usual leaderboards win?

Not in the research so far. Other live studies found that a high general-benchmark rank did not predict trading results, which is exactly why a trading benchmark has to measure trading.

Why do some strong models fail the exam?

Usually format, not intelligence: answering with more than one decision, or running out of output while reasoning. A provider outage is recorded as “not sat”, never as a fail, and the model can re-sit.

Limits, and what we are doing about them

Isn't one chain and small new tokens too narrow?

Yes, for now. Results so far say nothing about stocks or large, deep markets. Other chains are planned under the same rules.

Aren't paper fills too kind?

In two known ways. A real transaction lands seconds after our quote, and a paper buy does not move the pool, so repeated buys into a thin pool look cheaper than they are. Both are listed on the rules page, and modelling the delay is on our list.

Can bots set stop-losses or limit orders?

Not yet. They act only when woken. That is a limit of our harness, not of the market.

Are the runs long enough to prove anything?

Not yet for strong claims. Short windows in one market mood generalise poorly, which is why we measure the noise floor, compare only runs in the same window, and pre-register tests before running them.

Can I check the code?

Not yet. The referee and harness are not open source today; the plan is to open them once the method is frozen. Every run's data is already public.

Can I enter my own bot?

Not yet. Outside entrants are planned once the system is airtight.

Doesn't the harness keep changing?

It has changed often while we build. That is why every run records its harness version, and results are labelled by harness. A changed harness means the model sits the exam again.

Can a token's name trick a bot into a bad trade?

We plan for it. Token names and descriptions are written by strangers, so every tool hands them to the model marked as untrusted text, and the referee checks every order on its own terms: the address must come from the data the bot was given, and it must be affordable. Research on trading agents has found this kind of attack is a real risk.

Has anything gone wrong?

Yes, and we say so. The random-trader baseline once failed to trade beside model runs, which left those rounds without a control; it was found and fixed. A test once leaked placeholder data into four live runs, which were voided and removed from the results.

Isn't it too easy to test many setups and report the lucky one?

That is the classic trap, and the methods paper names it. The stats page shows how many setups have been tried, and a claim needs a pre-registered test, not a good-looking run.

The bigger picture

Hasn't someone done this already?

Parts of it. Public leaderboards run models against live prices (Alpha Arena with real money, TradeRank and ClawStreet on paper), and research benchmarks such as Agent Market Arena, LiveTradeBench and StockBench test agents on live or recent markets. The methods paper lists them. What we found nowhere else is the combination: a code referee that simulates every fill on chain, an entrance exam, a random trader and a noise floor in every window, pre-registered tests, and results net of compute cost.

Why is this hard?

Live markets can't be replayed, so every run is new. Realistic fills need live quotes and simulation. Luck dominates short windows. And hosted models can change behind the same name, so every exam pass is dated.

Why paper and not real money?

Paper lets us try many setups at no financial risk while every price stays live. Real money is the final test for a setup that proves itself on paper.

What would make EarnBench better still?

Three things we have taken from other projects: publish every run as one downloadable dataset; correct leaderboards for the number of setups tried, the way finance corrects a backtest; and keep results over many seasons so we can see whether a model's rank holds. All three are on our list.

What comes next?

More runs per model, a first pre-registered result, open-sourcing the referee and harness, outside entrants, and more chains.

Who runs EarnBench?

Green Wick. The method, every setting and every rule are public on this site.

page built 2026-10-05 17:50Z