EarnBench: A Controlled Paper-Trading Benchmark
for Language-Model Agents on Live On-Chain Markets

EarnBench

Green Wick · earnbench.wick.pics

Working draft v0.2 · 24 September 2026 · results pending

Abstract

We describe the method of EarnBench, a benchmark that asks which combinations of language model, agent harness, prompt and tools make profitable trading decisions under realistic execution. Agents trade a fixed paper stake on a live blockchain market; every order is filled by an independent referee at a live multi-source quote and, since this draft, is re-checked by simulating the route's own transaction so that token taxes and failing transactions are charged as they would be on chain. Before a model trades it must pass an entrance exam with no errors. Results are scored as the liquidation value of the book in the starting currency and are only compared across agents that traded the same window, against a random-trading baseline and a measured noise floor. This draft fixes the design, the variables and the known limitations before results are reported, so that later changes to the method are visible as dated revisions rather than silent re-interpretation.

Keywords: LLM agents, trading, benchmark methodology, paper trading, execution realism, pre-registration.

1Introduction

Public claims that an AI system “can trade” are rarely comparable: they differ in market, capital, costs, fill assumptions and evaluation window, and they are usually reported after the fact. Benchmarks in other domains became useful once the task, the scoring and the environment were fixed and published in advance [2, 3]. EarnBench applies that discipline to a narrow, measurable question: given the same market, the same starting stake and the same execution rules, does this agent stack end with more than it started with, by more than chance? Trading is done on paper so that many stacks can be tried at no financial risk, but every price is live, so the market cannot be replayed or memorised.

2Design

2.1The unit under test

The unit is a stack: model Ă— harness Ă— prompt Ă— tools. A harness is a versioned manifest that fixes the tool list, the step and output limits and the provider; a prompt is stored verbatim. Both are identified by SHA-256 hashes recorded with every run, so any result can be re-run on an identical stack and any change produces a new identity rather than an edited one.

2.2Market and referee

Agents trade on Robinhood Chain (chain id 4663), an EVM network with Uniswap-style pools [1]. Each run starts with US$100 of ETH, priced when the run opens (runs before the change on 2026-09-24 started with a fixed 0.036 ETH). An agent is woken on a fixed interval (10 minutes by default; some runs test other intervals, stated in the run's prompt), receives a data bundle and its own book, and may submit orders (buy, sell, open or close a concentrated-liquidity position). It never holds balances: a separate referee process does. For each order the referee takes a live quote for the order's exact size from its own router, and asks a board of rival aggregators only when its own route has no price (until 2026-09-24 21Z every order asked the full board); it fills at the best row net of gas, refusing boards older than 30 s, buys whose honeypot check is not clear, and trades whose gas exceeds their value. The chosen route's own transaction is then simulated from the paper wallet on the latest block (eth_simulateV1 with transfer tracing [4]; the wallet's balance slot is found through the node's access list [5]), and the agent is credited with what actually arrives, not with what the route reports. A transaction that reverts is not filled. Two ledgers are kept from the same boards: the headline ledger counts only rows proven by simulation; a second ledger counts any row. Holdings are marked by a full-size live sell quote right after every turn and on the clock every 30 minutes (a sixth of the window when that is shorter), and the run ends with a full live liquidation.

2.3Protocol

  1. Entrance exam. Three fixed trading turns captured from a real run, played under the exact limits of a live turn. A model passes only if every check holds on every turn: the provider answers; the answer arrives within the time, step and output limits; it is exactly one valid decision object; the model uses its tools at least once and every tool call is well-formed; it names the inputs it relied on; every address it orders came from its data; and every order is affordable against the book it was shown. A pass belongs to one model on one harness identity.
  2. Plan turn. Before its window opens, the stack gets one long research turn (up to 20 minutes, same tools, a higher quote allowance, no book yet) and writes a trading plan. The plan is frozen into the run's recorded prompt and published with it, so every later trade can be judged against it. The window clock starts only after the plan is written, so preparation never uses trading time.
  3. Solo shakedown. One 1-hour run per passing stack, to show it can trade end-to-end. Not used for ranking.
  4. Noise floor. Four copies of one stack trade the same window. The spread of their returns estimates how much chance alone moves a score.
  5. Comparison. Stacks trade the same window side by side, with a random trader on the same cadence and order sizes as a baseline.
  6. Evolution arena. A later format in which the weakest agent is periodically removed and the strongest copied with one deliberate change; described separately.

3Variables

Table 1: What is varied, what is measured, and what is held fixed.
Varied (the question)Measured (the answer)Held fixed (the control)
Model and its reasoning effortLiquidation return, headline ledger, in ETHChain, starting stake (US$100 of ETH at the open)
Harness: tools, step and output limitsEntrance-exam checks passedWake cadence (10 min), window per format
Prompt and prompting frameworkFills, refusals and their reasonsFill rule, quote sources, simulation of every fill
Tool and data accessCosts paid: gas, pool fees, tax, slippageMarking rule (after each turn and on the clock, full-size sell)
(Later) evolution settingsTokens used and wall time per turnStarting data bundle layout per harness

4Scoring and inference

The score of a run is the percentage change of the headline ledger's liquidation value, measured in ETH. Because every agent in a window experiences the same ETH price move, rankings within a window are identical in ETH or in dollars; the currency is a presentation choice. Holding ETH scores exactly zero and is therefore the natural floor. A difference between two stacks is reported as a finding only when (a) both traded the same window, (b) it exceeds the measured noise floor for that window length (two standard errors of the copies' spread), and (c) both beat the random trader. Every score is published with its n; a single 1-hour run is treated as an anecdote, not a result. With many stacks tested, some will look good by chance [6, 7]; the number of stacks tried is therefore published beside every leaderboard.

Every run is also reported after compute: the trading result minus what the model's own calls would cost at the paid list price for that model (nothing is charged during the benchmark, which uses free models or a subscription). Because profit and compute are in the same unit, this yields a single, sharper question than raw return: at what stake does a model's trading cover the cost of running it? Since compute per decision is roughly fixed while profit scales with the stake until liquidity binds, each stack has a break-even range of stakes, possibly empty.

Cost can be counted more than one way, and whether a model pays for itself depends on whose cost is counted. We intend to report five bases: API cost, the provider's paid list price for the tokens the run used; compute cost, the raw hardware time needed to serve those tokens at retail rental prices, without a large provider's economies of scale; cost to the lab, what serving them costs the lab itself at its own scale; our cost, what the run actually cost the operator: nothing on a free tier, or a share of a flat subscription by a stated rule; and environmental cost, the energy the run consumed and the emissions that follow from it. Only the first is observed directly, from published price lists, and it is the only basis reported so far. Labs do not publish the inputs to the others, so those are estimates built from public figures (hardware prices, published serving throughput, measured energy per token). Each estimated basis will be published with its method, its sources and a range rather than a single figure where the inputs are uncertain. It will be labelled as an estimate wherever it appears, and a ranking under an estimated basis is only as reliable as the estimate behind it.

5Threats to validity

  1. Execution delay. Fills occur at the quote taken when the order arrives; a real transaction lands seconds later. In an earlier live system of ours, live trades lagged their paper twins materially on identical signals, and the gap was execution. Not yet modelled; a delayed re-quote with a slippage bound is planned.
  2. Own price impact. A paper buy does not move the pool, so repeated buys into one thin pool look cheaper than they are. A single round trip is approximately unaffected only while its size is small against the pool's depth; in this chain's thinnest launch pools a full-size sell fills well below the mid.
  3. No orders between turns. Agents cannot place stop or limit orders; they act only when woken.
  4. Honeypot gate. The quote engine's honeypot check refuses some tokens as unknown; those tokens are untradeable here even where a real trader could buy them.
  5. Simulation coverage. A fill that cannot be simulated (lane outage, unusual token storage) is refused on the headline ledger for buys and tagged for sells, so that an exit is never trapped by our own infrastructure.
  6. Model drift and availability. Free hosted models change behind a stable name and are sometimes unavailable; exam passes are dated, and an outage is recorded as “not sat”, never as a failure.
  7. Regime and sample size. Short windows in one market regime generalise poorly; sequential solo runs face different markets and are not compared.
  8. Exam representativeness. The three exam turns come from one early run; a model can pass the exam and still trade poorly. The exam tests ability to play, not skill.
  9. Decision latency and turn order. Fills are priced when an order arrives, so a slower model trades at later prices. We count that as part of the model's result, as it would be in real trading, and publish each model's seconds per turn. What is not the model's doing is controlled: when several agents share a window their turn order rotates each tick, provider rate-limit waits are recorded, and a turn a provider refuses is labelled unavailable rather than counted against the model.
  10. Liquidity provision. Paper LP fees are the position's true share of real swaps in range and impermanent loss follows the real price path, but a paper position does not attract or deter flow.

6Related work

Live trading benchmarks. Several recent benchmarks evaluate language-model agents on live or post-cutoff markets rather than historical backtests: Agent Market Arena across crypto and stocks [8], AI-Trader across US stocks, A-shares and crypto [9], LiveTradeBench across US stocks and Polymarket prediction markets [10], and StockBench on a contamination-free multi-month stock window [11]. Their headline findings point the same way as our design: agent frameworks explain more of the outcome than the model backbone [8]; general ability, or a high chat-arena rank, does not imply trading skill [9, 10]; and most models struggle to beat buy-and-hold [11]. Earlier work evaluated agents on historical environments [12] or under adversarial market transforms [13]. Public leaderboards also run models against live prices, with real money [23] or on paper [24, 25]. In those we have examined, the model is the variable under a fixed prompt or an agent's own code, and we found no reported noise floor or random-trading baseline.

Trading agents. A large literature builds LLM trading agents with memory [14], tools and reflection [15, 16] and multi-agent roles such as analysts, debaters and risk managers [17], and reports large backtest gains. Those gains are measured on historical windows that may overlap training data, usually with light cost models; robustness and security studies of such schemes find frequent failures under small perturbations and attack [18, 19]. We treat these designs as candidate stacks to test, not as evidence that the stack earns.

Harnesses, selection and inference. The interface between model and environment can move results as much as the model [20], which is why the unit here is the whole stack. Tournament and evolutionary selection of trading agents appears in [21, 22]. For inference under many trials we follow the backtest-overfitting literature [6, 7].

What we add. Among the work above we found no benchmark that combines: an independent, non-model referee that simulates every fill on a live on-chain market; an entrance exam before any trading; comparison only within one window, against a random trader and a measured noise floor from identical copies; pre-registered tests; and results reported net of the model's compute cost. This survey was made on 2026-09-25 and will be extended.

7Status

At this draft, ten hosted chat models have sat the entrance exam. Two have passed, and both only on the harness whose output limit is 16,000 tokens rather than 6,000: reasoning models spend hidden reasoning inside that limit, and three models ran out of it before answering. The other common failures are answers that are not a single decision object and provider outages, which are recorded as not sat. Live counts are on the entrance page, the parameters on /param and the rules with their enforcement status on /rules. No performance result is reported in this version.

References

  1. H. Adams, N. Zinsmeister, M. Salem, R. Keefer, D. Robinson. Uniswap v3 Core. Whitepaper, 2021.
  2. P. Liang et al. Holistic Evaluation of Language Models. arXiv:2211.09110, 2022.
  3. C. E. Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv:2310.06770, 2023.
  4. Ethereum Execution APIs. eth_simulateV1 specification. github.com/ethereum/execution-apis.
  5. V. Buterin, M. Swende. EIP-2930: Optional access lists. Ethereum Improvement Proposals, 2020.
  6. D. H. Bailey, J. Borwein, M. LĂłpez de Prado, Q. J. Zhu. The Probability of Backtest Overfitting. Journal of Computational Finance, 2017.
  7. C. R. Harvey, Y. Liu, H. Zhu. …and the Cross-Section of Expected Returns. Review of Financial Studies 29(1), 2016.
  8. L. Qian et al. When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents. arXiv:2510.11695, 2025.
  9. T. Fan et al. AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets. arXiv:2512.10971, 2025.
  10. H. Yu, F. Li, J. You. LiveTradeBench: Seeking Real-World Alpha with Large Language Models. arXiv:2511.03628, 2025.
  11. Y. Chen et al. StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? arXiv:2510.02209, 2025.
  12. H. Li et al. INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent. arXiv:2412.18174, 2024.
  13. X. Yuan, H. Xu, S. Xu. TraderBench: How Robust Are AI Agents in Adversarial Capital Markets? arXiv:2603.00285, 2026.
  14. Y. Yu et al. FinMem: A Performance-Enhanced LLM Trading Agent with Layered Memory and Character Design. arXiv:2311.13743, 2023.
  15. W. Zhang et al. A Multimodal Foundation Agent for Financial Trading: Tool-Augmented, Diversified, and Generalist. arXiv:2402.18485, 2024.
  16. Y. Li et al. A Reflective LLM-based Agent to Guide Zero-shot Cryptocurrency Trading (CryptoTrade). arXiv:2407.09546, 2024.
  17. Y. Xiao et al. TradingAgents: Multi-Agents LLM Financial Trading Framework. arXiv:2412.20138, 2024.
  18. L. Yan et al. TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful? arXiv:2512.02261, 2025.
  19. M. Wang et al. SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes. arXiv:2609.19705, 2026.
  20. J. Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. arXiv:2405.15793, 2024.
  21. L. Zhao et al. ContestTrade: A Multi-Agent Trading System Based on Internal Contest Mechanism. arXiv:2508.00554, 2025.
  22. S. Kim et al. EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents. arXiv:2609.17632, 2026.
  23. Nof1. Alpha Arena. nof1.ai, 2025.
  24. TradeRank. LLM trading seasons. traderank.ai, 2026.
  25. ClawStreet. Agent trading arena. clawstreet.io, 2026.
Revision log

The PDF is v0.2 as first published. Each change is added here with its date; earlier versions are kept.