What EarnBench does, how it works, what people tend to misread, and every limit we know of, with what we are doing about it. 45 questions, short answers.
A benchmark for one question: can an AI trading setup make money in a real market, by more than luck? We test the whole stack (model, agent harness, prompt and tools), not the model alone.
Yes. The money is paper, but every price is live. Each order fills at a live quote for exactly its size, taken at that moment (our own route first, rival aggregators when ours has no price), so the market can't be replayed or memorised.
We work hard to stop that. The route's own transaction is simulated on the latest block, so token taxes and transactions that would fail are charged as they would be on chain, and the headline score counts only fills on routes checked on chain. A second line, which counts every fill, is published beside it.
No. A model never holds a balance. A separate referee keeps the books, fills orders and values holdings; the model can only send orders.
By liquidation value: every holding is valued with a live, full-size sell quote, not the mid price. That is what the bot could actually walk away with. Holding ETH scores exactly zero.
Every run sits beside a random trader in the same window. We also run copies of one model side by side to measure how far luck alone moves a score. A difference only counts once it clears that noise, and a single 1-hour run is treated as an anecdote, not a result.
Every harness and prompt is identified by a SHA-256 hash, recorded with each run. Registered tests are frozen by those hashes before they start, and any change is written down as a dated deviation. The methods paper keeps a dated revision log.
So that a bad result means a bad decision, not a broken answer. Before it trades, a model must play three real trading turns without a single error: a valid answer within the limits, real tool use, and orders that are affordable and use only addresses it was shown. Results are on the entrance exam page.
Yes. Gas, pool fees, token taxes and slippage are all paid. Each run is also reported after compute: the trading result minus what the model's own calls would cost at its paid list price.
Most contests give each model the same prompt and rank the returns. We test the whole setup (model, harness, prompt and tools), a code referee rather than the model fills and checks every trade, comparisons are made within one window against a random trader and copies of the same model, tests are registered before they run, and results are shown after the cost of the model's own calls. The methods paper compares us with the closest projects.
Because the question is whether it makes money. A judge model can be persuaded by a confident explanation of a losing trade; a referee that only counts what the wallet would really hold cannot.
No. It trades prices that did not exist when it was trained, so there is no past answer to remember. That is a common weakness of backtests on historical data, and a live market avoids it.
Robinhood Chain (chain id 4663), a public EVM network with Uniswap-style pools and many newly launched tokens. EarnBench is not affiliated with Robinhood.
US$100 of ETH, priced when its run opens.
It is woken on a fixed schedule (every 10 minutes by default). Each time it sees its data and its own book and may send up to five orders: buy, sell, or open or close a liquidity position.
Whatever its harness gives it, and nothing else. Each harness's exact tools and limits are on the tools page.
Yes. Before its window opens it gets one research turn of up to 20 minutes to write a trading plan. The plan is published with the run, so every later trade can be judged against it.
1-hour runs prove a model can trade end to end and let us try variations quickly. Longer runs, such as 4 hours, are used for comparisons.
In current trials, a bot is told that taking no action disqualifies it. A run that then places no trade at all is left out of the results.
Models we can run without per-call charges: free tiers from several providers, and Claude on a subscription. The entrance exam page lists every model that has sat the exam and the stats page counts them.
An anonymous preview model on OpenRouter, free while its lab keeps it unnamed. If the lab later confirms it was the same model all along, its results will be renamed to the real model. If the released model is different, they stay under the preview name.
No. Every trade is on paper, priced and checked against the live market.
No. It means that setup did well in those windows on this market. Past windows promise nothing about future ones, and paper fills are kinder than real ones in two known ways (see Limits).
The entrance exam and stats pages show each model's best run, always with how many runs it had and their average beside it. Registered tests use every run, not the best one.
Every bot in a window sees the same ETH price move, so the ranking is the same either way. In ETH, doing nothing scores exactly zero, which makes the baseline obvious.
No. They share a market, not an order book. A paper trade does not move the pool, so one bot cannot change another's prices.
Not in the research so far. Other live studies found that a high general-benchmark rank did not predict trading results, which is exactly why a trading benchmark has to measure trading.
Usually format, not intelligence: answering with more than one decision, or running out of output while reasoning. A provider outage is recorded as “not sat”, never as a fail, and the model can re-sit.
Yes, for now. Results so far say nothing about stocks or large, deep markets. Other chains are planned under the same rules.
In two known ways. A real transaction lands seconds after our quote, and a paper buy does not move the pool, so repeated buys into a thin pool look cheaper than they are. Both are listed on the rules page, and modelling the delay is on our list.
Not yet. They act only when woken. That is a limit of our harness, not of the market.
Not yet for strong claims. Short windows in one market mood generalise poorly, which is why we measure the noise floor, compare only runs in the same window, and pre-register tests before running them.
Not yet. The referee and harness are not open source today; the plan is to open them once the method is frozen. Every run's data is already public.
Not yet. Outside entrants are planned once the system is airtight.
It has changed often while we build. That is why every run records its harness version, and results are labelled by harness. A changed harness means the model sits the exam again.
We plan for it. Token names and descriptions are written by strangers, so every tool hands them to the model marked as untrusted text, and the referee checks every order on its own terms: the address must come from the data the bot was given, and it must be affordable. Research on trading agents has found this kind of attack is a real risk.
Yes, and we say so. The random-trader baseline once failed to trade beside model runs, which left those rounds without a control; it was found and fixed. A test once leaked placeholder data into four live runs, which were voided and removed from the results.
That is the classic trap, and the methods paper names it. The stats page shows how many setups have been tried, and a claim needs a pre-registered test, not a good-looking run.
Parts of it. Public leaderboards run models against live prices (Alpha Arena with real money, TradeRank and ClawStreet on paper), and research benchmarks such as Agent Market Arena, LiveTradeBench and StockBench test agents on live or recent markets. The methods paper lists them. What we found nowhere else is the combination: a code referee that simulates every fill on chain, an entrance exam, a random trader and a noise floor in every window, pre-registered tests, and results net of compute cost.
Live markets can't be replayed, so every run is new. Realistic fills need live quotes and simulation. Luck dominates short windows. And hosted models can change behind the same name, so every exam pass is dated.
Paper lets us try many setups at no financial risk while every price stays live. Real money is the final test for a setup that proves itself on paper.
Three things we have taken from other projects: publish every run as one downloadable dataset; correct leaderboards for the number of setups tried, the way finance corrects a backtest; and keep results over many seasons so we can see whether a model's rank holds. All three are on our list.
More runs per model, a first pre-registered result, open-sourcing the referee and harness, outside entrants, and more chains.
Green Wick. The method, every setting and every rule are public on this site.